CAUSAL Lab | Zhiheng Zhang(张智恒)| SSDS, SUFE

欢迎访问张智恒的主页!本实验室名为CAUSAL (Causal Analysis for Underlying Structures And Learning),致力于发展弱假设和复杂环境下的高效通用自动化的因果推断方法,聚焦于潜在结构的刻画,并将其与现代学习算法和决策系统融合。请查看实验室因果推断入门指南

张智恒自2025年8月起担任上海财经大学统计与数据科学学院常任轨助理教授,他同时兼职隶属于上海财经大学大数据研究院。此前,他于清华大学交叉信息研究院(IIIS)获得博士学位,博士生导师为王禹皓教授。张智恒担任2025年度CCF-滴滴盖亚联合科研基金项目“统一的多treatment LTV因果模型”的联合负责人, AAAI 2025 人工智能会议 Causal Techniques 方向(AICT track)领域主席等。

他正在寻找博士生/硕士生/实习生。 对于想申请的同学, 请完成这个问卷。他的邮箱是zhangzhiheng@mail.shufe.edu.cn。此外,欢迎对因果推断感兴趣的老师和同学参加我们的论文讨论会,可通过微信公众号CAUSAL-lab-SSDS-SUFE加入。

📢 招生信息 / Recruitment

我正在寻找若干 博士生和硕士生。同时也长期招收科研实习生(支持远程线上实习)

我对博士生的基本期望是:

  • 具备良好的品德和沟通能力,内驱力强,作风扎实。喜欢科研探索(鼓励自主探索感兴趣的方向。我会在能力范围内尽力指导;在能力范围外建立合作);
  • 满足以下两项中的任意一项:扎实的数学基础(尤其是概率统计,专业不限);出色的编程能力;
  • 将来有意愿走上科研道路。

我对硕士生的基本期望是:

  • 具备良好的品德和沟通能力,作风扎实;
  • 对于有志于学界的,我会跟博士生共同带领你进行科研,并支持你的继续深造;对于有志于业界的,我希望你主动求职的是算法科研岗位,我会根据适配度将你推荐到相关方向的前沿研究部门实习。此外,至少在在读期间与我完成一个你自己的科研课题,作为你的毕业论文。

如果你有兴趣,欢迎随时与我联系。其中,如果你想读博,请至少提前半年联系我;如果你想读硕,请至少提前两个月联系我。我们会通过提前合作来判断是否彼此合适。

学术体系图

为理解并构建因果学习系统的核心理论结构,实验室的长期目标可概括为以下三个基础问题:

  1. 观测性研究中:如何系统刻画“模型假设—观测数据—可识别边界”之间的传导机制,从而揭示各种因果假设对可识别性的根本影响?

  2. 实验设计与推断中:如何定量描述并评估“应用场景属性—实验设计与算法结构—统计效率”之间的性能极限,并基于此开发具有最优(或近最优)性质的设计与推断框架?

  3. 离线/在线学习与决策中:如何从数学结构上统一机器学习、经济管理与统计推断等不同领域的优化目标,揭示它们之间的基本兼容性与最优可达边界?

为回答上述问题,实验室的总体研究路线遵循由理论到方法、由方法到实践的逐层递进结构:

  1. 从基础假设的违背出发,构建更具包容性的因果推断框架,例如:探索 unconfoundedness、overlap、SUTVA 等假设的违背情形;

  2. 将这些基础结构与现代统计与机器学习方法融合,发展更高效、更稳定、更具可扩展性的识别与估计技术,例如:最优传输、代理变量与负控方法、共型预测、minimax优化、在线学习等;

  3. 进一步将理论与方法扩展至带有现实约束的任务设置,例如:输入/输出复杂结构、小样本学习、动态/缺失网络结构等;

  4. 最终形成能向现实场景有效辐射的因果推断体系,服务于社交网络分析、博弈论环境、优化决策,并落地于推荐系统、派单机制、市场干预策略、大模型行为建模等;

其中:1 关注输入层面更宽松的结构假设,2 聚焦算法层面精准且高效的识别—估计机制,3-4 主要面向输出层面复杂而贴近现实的决策与推断任务。

为逐步实现,目前实验室正主要围绕以下具体方向做研究:

  1. 理论基础:对偶理论(例如最优传输)驱动的因果识别与评估;

  2. 工业落地:面向工业现实约束的因果推断方法论(如RCT&OBS, pre&opt, structural data types,网络结构下的在线实验设计与推断);

  3. 智能拓展:LLM Reward Model, Causal Tabular Foundation Model, Causality in Stochastic Algorithms。

  4. 因果整合:从实验设计到科学发现——构建因果驱动的可信 AI 决策系统。 面对高维 treatment、网络暴露、组合行动空间和复杂行为轨迹,单个实验已难以充分刻画真实系统中的干预机制,未来的核心挑战也不再只是设计更复杂的实验,而是如何把分散的实验对象、实验结果与上下文条件组织成可计算、可比较、可桥接的因果结构,从而支撑长期可解释、可评估、可控制的 AI 决策系统。为此,我们将借助大模型的语义表征与生成式能力,首先构建从实验档案到科学发现的因果图谱层,将既有实验中的 treatment、outcome、context 与 effect 统一映射到结构化因果空间,研究多个实验何时可以组合、何时存在结论冲突,何处只能得到部分识别边界,并进一步通过最优 bridge experiments 主动缩小不可识别区域;在此基础上,面向多智能体互动、平台机制、市场干预与大模型行为系统,我们将设计因果伪奖励、遗憾校准分配、动态网络控制和战略因果 AI 机制,使复杂系统不仅具有良好的短期表现,而且能够在长期演化中保持可靠性;进一步地,我们还将通过稀疏 LLM 比较、低秩任务—模型能力结构与不确定性认证,构建可解释的因果路由器,使 AI 系统能够判断不同模型和策略在何种任务、何种环境下具有真正可迁移、可部署的因果优势。由此,整个研究形成一个从历史实验整理、因果图谱构建、识别冲突发现、桥接实验设计、在线决策优化到新证据反哺图谱的闭环,最终将研究重点从“如何设计复杂实验”推进到“如何组织实验空间、识别机制边界,构建可信决策机制”,形成面向科学发现的可识别学习理论。

ATLAS计划

他欢迎相关的交流或合作。 若感兴趣,可邮件联系或添加他的WeChat

Welcome to Zhiheng’s Lab! This lab is called CAUSAL (Causal Analysis for Underlying Structures And Learning), focusing on developing principled methods for causal analysis by uncovering underlying structures and integrating them with modern learning and decision-making systems. Please read the onboarding guide of CAUSAL lab.

Zhiheng Zhang (pronunciation: Zhee-hung Jahng) is a tenure-track Assistant Professor in the School of Statistics and Data Science at the Shanghai University of Finance and Economics, starting from 2025.08. He is also affiliated with the Institute of Big Data Research. Previously, he received his PhD from the Institute for Interdisciplinary Information Sciences at Tsinghua University, advised by Professor Yuhao Wang. Zhiheng Zhang is the co-Principal Investigator of the 2025 CCF-DiDi Gaia Collaborative Research Fund project titled “Unified multi treatment LTV causal model” and the area chair of AAAI2025 AICT track. He is looking for PhD/Master/Interns. His email is zhangzhiheng@mail.shufe.edu.cn.

Research Statement

To understand and develop the core theoretical structure of causal learning systems, the lab’s long-term agenda centers on the following three foundational questions:

  1. In observational studies: How can we systematically characterize the mechanism linking model assumptions, observed data, and identification boundaries, thereby revealing the fundamental effects of different causal assumptions on identifiability?

  2. In experimental design and inference: How can we quantitatively characterize and evaluate the performance limits arising from the interplay among application-setting attributes, experimental designs and algorithmic structures, and statistical efficiency, and on this basis develop design and inference frameworks with optimal (or near-optimal) properties?

  3. In offline and online learning and decision-making: How can we mathematically unify the optimization objectives of machine learning, economics and management, and statistical inference, and reveal their fundamental compatibility and optimal attainable frontiers?

To answer these questions, the lab follows a layered research trajectory that proceeds from theory to methodology and from methodology to practice:

  1. Begin with violations of foundational assumptions and construct more inclusive causal-inference frameworks—for example, by studying violations of unconfoundedness, overlap, and SUTVA;

  2. Integrate these foundational structures with modern statistical and machine-learning methods to develop more efficient, stable, and scalable techniques for identification and estimation, including optimal transport, proxy-variable and negative-control methods, conformal prediction, minimax optimization, and online learning;

  3. Extend the resulting theory and methodology to tasks with realistic constraints, such as complex input or output structures, small-sample learning, and dynamic or partially observed network structures;

  4. Ultimately establish a causal-inference framework that can translate effectively into real-world settings, supporting social-network analysis, game-theoretic environments, and optimization-based decision-making, with applications to recommender systems, dispatch mechanisms, market interventions, and the behavioral modeling of large models.

In this progression, Step 1 focuses on more flexible structural assumptions at the input level; Step 2 develops precise and efficient identification-estimation mechanisms at the algorithmic level; and Steps 3-4 address complex, realistic decision and inference tasks at the output level.

To advance this agenda, the lab is currently pursuing four concrete research directions:

  1. Theoretical foundations: Causal identification and evaluation driven by duality theory, including optimal transport;

  2. Industrial deployment: Causal-inference methodologies tailored to real-world industrial constraints, including RCT & OBS, pre & opt, structural data types, and online experimental design and inference under network structures;

  3. Intelligent extensions: LLM reward models, causal tabular foundation models, and causality in stochastic algorithms;

  4. Causal integration: From experimental design to scientific discovery—building causality-driven trustworthy AI decision systems. With high-dimensional treatments, network exposure, combinatorial action spaces, and complex behavioral trajectories, a single experiment can no longer adequately characterize intervention mechanisms in real systems. The central challenge is therefore no longer merely to design more complex experiments, but to organize dispersed experimental entities, results, and contextual conditions into causal structures that are computable, comparable, and bridgeable, thereby supporting AI decision systems that remain interpretable, evaluable, and controllable over the long term. To this end, we will harness the semantic representations and generative capabilities of large models to first construct a causal-atlas layer connecting experimental archives to scientific discovery. This layer will map the treatments, outcomes, contexts, and effects from existing experiments into a unified structured causal space; determine when findings from multiple experiments can be combined, when their conclusions conflict, and where only partial-identification bounds are available; and actively shrink unidentified regions through optimal bridge experiments. Building on this foundation, we will develop causal pseudo-rewards, regret-calibrated allocation, dynamic network control, and strategic causal-AI mechanisms for multi-agent interactions, platform mechanisms, market interventions, and large-model behavioral systems, so that complex systems achieve not only strong short-term performance but also reliability under long-term evolution. We will further construct interpretable causal routers through sparse LLM comparisons, low-rank task-model capability structures, and uncertainty certification, enabling AI systems to determine which models and strategies possess genuinely transferable and deployable causal advantages for a given task and environment. Together, these components form a closed loop—from organizing historical experiments, constructing causal atlases, and detecting identification conflicts, through designing bridge experiments and optimizing online decisions, to feeding new evidence back into the atlas. The goal is to shift the research frontier from “how to design complex experiments” to “how to organize the experimental space, identify mechanism boundaries, and build trustworthy decision mechanisms,” thereby developing a theory of identifiable learning for scientific discovery.

Open to discussions and collaborations at any time. You can send him an email or add his WeChat.

Selected Honours

2018 Chinese Undergraduate Mathematics Competition Final (CMC), Gold Medal (Top 10 in China)

Academic Service

Reviewer: ICML, NeurIPS, ICLR, AISTATS, UAI, AAAI, ACM Transactions on Information Systems (TOIS), Transactions of Mobile Computing (TMC), Journal of Machine Research (JMLR), Journal of the Royal Statistical Society, Series B (JRSSB), Journal of the American Statistical Association (JASA), Electronic Journal of Statistics (EJS), Journal of Computational and Graphical Statistics (JCGS)

Area Chair: AAAI 2025, Artificial Intelligence with Causal Techniques (AICT) track

Session Chair: Causality and Machine Learning, Tsinghua Sanya International Mathematics Forum (TSIMF) 2026

中国现场统计研究会因果推断分会理事