CAUSAL Lab | Zhiheng Zhang(张智恒)| SSDS, SUFE

语言 / Language

欢迎访问张智恒的主页!本实验室名为 CAUSALCausal Analysis for Underlying Structures And Learning),致力于在弱假设和复杂环境下发展高效、通用、自动化的因果推断方法,聚焦潜在结构的刻画,并将其与现代学习算法和决策系统相融合。请查看 CAUSAL 实验室因果推断入门指南

张智恒(英文姓名读音近似 “Zhee-hung Jahng”)自 2025 年 8 月起任上海财经大学统计与数据科学学院常任轨助理教授,同时兼职隶属于上海财经大学大数据研究院。此前,他于清华大学交叉信息研究院(IIIS)获得博士学位,博士生导师为王禹皓教授。他担任 2025 年度 CCF—滴滴盖亚联合科研基金项目“统一的多处理长期价值(LTV)因果模型”联合负责人,并担任 AAAI 2025 人工智能与因果技术方向(AICT track)领域主席。

他正在招收博士生、硕士生和科研实习生。有意申请者请完成此问卷,或发送邮件至 zhangzhiheng@mail.shufe.edu.cn。此外,欢迎对因果推断感兴趣的老师和同学参加实验室的论文讨论会;可通过微信公众号 CAUSAL-lab-SSDS-SUFE加入。

📢 招生信息

我正在招收若干博士生和硕士生,同时长期招收科研实习生(支持远程线上实习)

我对博士生的基本期望是:

  1. 具备良好的品德与沟通能力,内驱力强,作风扎实,热爱科研探索。鼓励自主探索感兴趣的方向:在我的能力范围内,我会尽力指导;超出我能力范围的方向,我会协助建立合作;
  2. 至少具备以下两项能力之一:扎实的数学基础(尤其是概率统计,专业不限);出色的编程能力;
  3. 有志于未来从事科研工作。

我对硕士生的基本期望是:

  1. 具备良好的品德与沟通能力,作风扎实;
  2. 对于有志于学界的学生,我会与博士生共同指导其开展科研,并支持其继续深造;对于有志于业界的学生,我希望其主动求职目标是算法科研岗位,并会根据适配度推荐其前往相关方向的前沿研究部门实习。此外,每位硕士生至少应在读期间与我完成一项由自己主导的科研课题,并以此作为毕业论文。

如有兴趣,欢迎随时与我联系。计划申请博士者请至少提前半年联系;计划申请硕士者请至少提前两个月联系。我们将通过前期合作判断彼此是否合适。

学术体系图

研究陈述

为理解并构建因果学习系统的核心理论结构,实验室的长期目标可概括为以下三个基础问题:

  1. 观测性研究: 如何系统刻画“模型假设—观测数据—可识别边界”之间的传导机制,从而揭示不同因果假设对可识别性的根本影响?
  2. 实验设计与推断: 如何定量刻画并评估“应用场景属性—实验设计与算法结构—统计效率”之间的性能极限,并据此开发具有最优(或近似最优)性质的设计与推断框架?
  3. 离线与在线学习及决策: 如何从数学结构上统一机器学习、经济与管理、统计推断等领域的优化目标,并揭示它们之间的基本兼容性与最优可达边界?

为回答上述问题,实验室的总体研究路线遵循从理论到方法、再从方法到实践的逐层递进结构:

  1. 从基础假设的违背出发,构建包容性更强的因果推断框架,例如研究无混杂性(unconfoundedness)、重叠性(overlap)和稳定单元处理值假设(SUTVA)遭到违背的情形;
  2. 将这些基础结构与现代统计和机器学习方法相融合,发展更高效、更稳定、更具可扩展性的识别与估计技术,例如最优传输、代理变量与负控方法、共形预测、极小极大优化和在线学习等;
  3. 进一步将相关理论与方法扩展到具有现实约束的任务设定,例如复杂的输入或输出结构、小样本学习,以及动态或部分观测的网络结构;
  4. 最终形成能够有效辐射现实场景的因果推断体系,服务于社交网络分析、博弈论环境和基于优化的决策,并应用于推荐系统、派单机制、市场干预策略和大模型行为建模。

在这一递进路线中,第 1 步关注输入层面更宽松的结构假设;第 2 步聚焦算法层面精准且高效的识别—估计机制;第 3—4 步主要面向输出层面复杂且贴近现实的决策与推断任务。

为逐步推进上述目标,目前实验室主要围绕以下四个具体方向开展研究:

  1. 因果理论评估: 面向基于设计的推断(design-based inference)开展 Lean 形式化证明,维护开放问题,并基于对偶理论筛选和评估算法;
  2. 结构规律发现:从实验设计走向科学发现——构建因果驱动的可信 AI 决策系统。 面对高维处理、网络暴露、组合行动空间和复杂行为轨迹,单个实验已难以充分刻画真实系统中的干预机制。未来的核心挑战不再只是设计更复杂的实验,而是如何将分散的实验对象、实验结果与情境条件组织为可计算、可比较、可桥接的因果结构,从而支撑长期可解释、可评估、可控制的 AI 决策系统。为此,我们将借助大模型的语义表征与生成能力,首先构建连接实验档案与科学发现的因果图谱层,将既有实验中的处理(treatment)、结果(outcome)、情境(context)与效应(effect)统一映射到结构化因果空间,研究多个实验何时可以组合、何时存在结论冲突、何处只能得到部分识别边界,并进一步通过最优桥接实验(bridge experiments)主动缩小不可识别区域。在此基础上,面向多智能体互动、平台机制、市场干预与大模型行为系统,我们将设计因果伪奖励、遗憾校准分配、动态网络控制和战略因果 AI 机制,使复杂系统不仅具有良好的短期表现,而且能在长期演化中保持可靠性。进一步地,我们还将通过稀疏 LLM 比较、低秩任务—模型能力结构与不确定性认证,构建可解释的因果路由器,使 AI 系统能够判断不同模型和策略在何种任务、何种环境下具有真正可迁移、可部署的因果优势。由此,整个研究形成一个闭环:从整理历史实验、构建因果图谱、发现识别冲突,到设计桥接实验、优化在线决策,再由新证据反哺图谱。最终,研究重点将从“如何设计复杂实验”推进到“如何组织实验空间、识别机制边界并构建可信决策机制”,从而形成面向科学发现的可识别学习理论;
  3. 智能估计与决策: 大语言模型奖励建模(LLM reward modeling)、因果表格数据基础模型(foundation models for causal tabular data),以及随机算法中的因果性;
  4. 面向产业的交叉应用: 针对产业现实约束的因果推断方法论,包括 RCT & OBS、pre & opt、结构化数据类型,以及网络结构下的在线实验设计与推断。

ATLAS 计划

他始终欢迎相关交流与合作。如有兴趣,可发送邮件或添加他的 WeChat

代表性荣誉

2018 年全国大学生数学竞赛决赛(CMC)金牌(全国前 10 名)

学术服务

  • 审稿人: ICML、NeurIPS、ICLR、AISTATS、UAI、AAAI、ACM Transactions on Information Systems(TOIS)、IEEE Transactions on Mobile Computing(TMC)、Journal of Machine Learning Research(JMLR)、Journal of the Royal Statistical Society: Series B(JRSSB)、Journal of the American Statistical Association(JASA)、Electronic Journal of Statistics(EJS)、Journal of Computational and Graphical Statistics(JCGS)
  • 领域主席: AAAI 2025,Artificial Intelligence with Causal Techniques(AICT)方向
  • 分会场主席: Causality and Machine Learning,Tsinghua Sanya International Mathematics Forum(TSIMF)2026
  • 理事: 中国现场统计研究会因果推断分会

Welcome to Zhiheng’s homepage! The lab is named CAUSAL (Causal Analysis for Underlying Structures And Learning). We develop efficient, general-purpose, and automated methods for causal inference under weak assumptions and in complex environments, with an emphasis on characterizing underlying structures and integrating them with modern learning algorithms and decision-making systems. Read the CAUSAL Lab Onboarding Guide to Causal Inference.

Since August 2025, Zhiheng Zhang (pronounced “Zhee-hung Jahng”) has been a tenure-track Assistant Professor at the Shanghai University of Finance and Economics, with his appointment in the School of Statistics and Data Science. He is also affiliated with the Institute of Big Data Research. Previously, he received his PhD from the Institute for Interdisciplinary Information Sciences (IIIS) at Tsinghua University, where he was advised by Professor Yuhao Wang. He is a co-Principal Investigator of the 2025 CCF–DiDi Gaia Collaborative Research Fund project, “A Unified Causal Model for Multi-Treatment Lifetime Value (LTV),” and served as an Area Chair for the Artificial Intelligence with Causal Techniques (AICT) track at AAAI 2025.

He is recruiting PhD students, master’s students, and research interns. Prospective applicants should complete this questionnaire or email zhangzhiheng@mail.shufe.edu.cn. Faculty members and students interested in causal inference are also welcome to join the lab’s paper-reading group through the CAUSAL-lab-SSDS-SUFE WeChat official account.

📢 Recruitment

I am recruiting several PhD and master’s students and also welcome research interns on an ongoing basis (remote internships are available).

What I expect from PhD students:

  1. Integrity, strong communication skills, self-motivation, diligence, and a genuine enthusiasm for research. Students are encouraged to explore topics that interest them independently: I will provide the best guidance I can within my areas of expertise and help establish collaborations when a topic extends beyond them;
  2. At least one of the following: a solid mathematical foundation, particularly in probability and statistics, regardless of undergraduate major; or excellent programming skills;
  3. An aspiration to pursue a research career.

What I expect from master’s students:

  1. Integrity, strong communication skills, and diligence;
  2. For students planning an academic career, I will mentor their research jointly with PhD students and support their pursuit of further study. For students planning an industry career, I expect them to target research-oriented algorithm roles proactively and, when there is a good fit, will recommend them for internships in leading research groups in relevant areas. In addition, each master’s student should complete at least one self-directed research project with me during the degree and develop it into their thesis.

If you are interested, please feel free to contact me. Prospective PhD applicants should reach out at least six months in advance, and prospective master’s applicants at least two months in advance. We will use preliminary collaboration to determine whether we are a good mutual fit.

Research roadmap

Research Statement

To understand and develop the core theoretical structure of causal learning systems, the lab’s long-term agenda centers on the following three foundational questions:

  1. Observational studies: How can we systematically characterize the mechanism linking model assumptions, observed data, and identification boundaries, thereby revealing the fundamental effects of different causal assumptions on identifiability?
  2. Experimental design and inference: How can we quantitatively characterize and evaluate the performance limits arising from the interplay among application-setting attributes, experimental designs and algorithmic structures, and statistical efficiency, and on this basis develop design and inference frameworks with optimal or near-optimal properties?
  3. Offline and online learning and decision-making: How can we mathematically unify the optimization objectives of machine learning, economics and management, and statistical inference, and reveal their fundamental compatibility and optimal attainable frontiers?

To answer these questions, the lab follows a layered research trajectory that proceeds from theory to methodology and from methodology to practice:

  1. Begin with violations of foundational assumptions and construct more inclusive causal-inference frameworks—for example, by studying violations of unconfoundedness, overlap, and the stable unit treatment value assumption (SUTVA);
  2. Integrate these foundational structures with modern statistical and machine-learning methods to develop more efficient, stable, and scalable techniques for identification and estimation, including optimal transport, proxy-variable and negative-control methods, conformal prediction, minimax optimization, and online learning;
  3. Extend the resulting theory and methodology to tasks with realistic constraints, such as complex input or output structures, small-sample learning, and dynamic or partially observed network structures;
  4. Ultimately establish a causal-inference framework that translates effectively into real-world settings, supporting social-network analysis, game-theoretic environments, and optimization-based decision-making, with applications to recommender systems, dispatch mechanisms, market interventions, and the behavioral modeling of large models.

In this progression, Step 1 focuses on more flexible structural assumptions at the input level; Step 2 develops precise and efficient identification–estimation mechanisms at the algorithmic level; and Steps 3–4 address complex, realistic decision and inference tasks at the output level.

To advance this agenda, the lab is currently pursuing four concrete research directions:

  1. Causal theory evaluation: Lean formalization for design-based inference, maintenance of open problems, and algorithm screening and evaluation based on duality theory;
  2. Discovery of structural regularities: From experimental design to scientific discovery—building causality-driven trustworthy AI decision systems. Given high-dimensional treatments, network exposure, combinatorial action spaces, and complex behavioral trajectories, no single experiment can adequately characterize intervention mechanisms in real-world systems. The central challenge ahead is therefore no longer simply to design more complex experiments, but to organize dispersed experimental objects, findings, and contextual conditions into causal structures that can be computed, compared, and bridged, thereby supporting AI decision systems that remain interpretable, evaluable, and controllable over the long term. To this end, we will leverage the semantic representations and generative capabilities of large models to first build a causal-atlas layer linking experimental archives to scientific discovery. This layer will map the treatments, outcomes, contexts, and effects from prior experiments into a unified, structured causal space; determine when findings across experiments can be combined, when conclusions conflict, and where only partial-identification bounds are available; and use optimal bridge experiments to actively shrink unidentified regions. Building on this foundation, we will develop causal pseudo-rewards, regret-calibrated allocation, dynamic network control, and strategic causal-AI mechanisms for multi-agent interactions, platform mechanisms, market interventions, and large-model behavioral systems, enabling complex systems not only to perform well in the short term but also to remain reliable as they evolve over time. We will further construct interpretable causal routers through sparse LLM comparisons, low-rank task–model capability structures, and uncertainty certification, allowing AI systems to determine which models and strategies possess genuinely transferable and deployable causal advantages in which tasks and environments. Together, these components form a closed loop: from organizing historical experiments, building causal atlases, and detecting identification conflicts, through designing bridge experiments and optimizing online decisions, to feeding new evidence back into the atlas. The goal is to shift the research focus from “how to design complex experiments” to “how to organize the experimental space, identify mechanism boundaries, and build trustworthy decision mechanisms,” thereby establishing a theory of identifiable learning for scientific discovery;
  3. Intelligent estimation and decision-making: LLM reward modeling, foundation models for causal tabular data, and causality in stochastic algorithms;
  4. Industry-facing interdisciplinary applications: Causal-inference methodologies tailored to real-world industrial constraints, including RCT & OBS, pre & opt, structural data types, and online experimental design and inference under network structures.

ATLAS initiative

He always welcomes discussions and collaborations in related areas. If you are interested, please email him or add him on WeChat.

Selected Honors

Gold Medal, 2018 Chinese Undergraduate Mathematics Competition Final (CMC; Top 10 nationwide)

Academic Service

  • Reviewer: ICML, NeurIPS, ICLR, AISTATS, UAI, AAAI, ACM Transactions on Information Systems (TOIS), IEEE Transactions on Mobile Computing (TMC), Journal of Machine Learning Research (JMLR), Journal of the Royal Statistical Society: Series B (JRSSB), Journal of the American Statistical Association (JASA), Electronic Journal of Statistics (EJS), and Journal of Computational and Graphical Statistics (JCGS)
  • Area Chair: AAAI 2025, Artificial Intelligence with Causal Techniques (AICT) track
  • Session Chair: Causality and Machine Learning, Tsinghua Sanya International Mathematics Forum (TSIMF) 2026
  • Council Member: Causal Inference Branch, Chinese Association for Applied Statistics (CAAS)