Selected Publications
(✉︎) Corresponding author; (‡) Equal contribution; (α-β order) alphabetical authorship ordering. You can also find my articles on my Google Scholar profile.
Journal Papers
- Partial Identification with Proxy of Latent Confoundings via Sum-of-ratios Fractional Programming: Extended
TL;DR: Proxy-based negative-control methods often require hard-to-verify invertibility or completeness; PI-SFP instead uses partial knowledge of P(W|U) to compute valid causal bounds with a globally convergent branch-and-bound algorithm. - Adjusting Auxiliary Variables Under Approximate Neighborhood Interference
TL;DR: Network-interference analyses often rely on restrictive outcome models; this paper provides a design-based adjustment that balances network covariates and improves precision under approximate neighborhood interference without requiring a correct outcome model. - A Minimax Learning Approach for Causal Inference Under Unmeasured Confounding with Negative Controls
TL;DR: Earlier negative-control methods relied on completeness, unique bridge functions, or parametric models; this paper removes the completeness requirement, largely relaxes uniqueness, and learns flexible bridges through minimax objectives with finite-sample guarantees.
Conference Papers
- Wasserstein Policy Learning for Distributional Outcomes
(Top conference in learning theory).TL;DR: Conventional offline policy learning reduces welfare to a scalar mean; this paper optimizes utilities of Wasserstein barycenters for distribution-valued outcomes and derives sharp regret bounds tied to policy-class complexity. - Causal Representation Learning with Optimal Compression and Complex Treatments
TL;DR: Multi-treatment ITE methods require heuristic tuning and scale poorly as treatments multiply; this paper derives an optimal balancing weight, introduces O(1)-scalable treatment aggregation, and adds a generative extension that preserves the treatment manifold's Wasserstein geometry. - Causal Matrix Completion under Multiple Treatments via Mixed Synthetic Nearest Neighbors
TL;DR: Synthetic Nearest Neighbors struggles when each treatment lacks enough fully observed anchors; MSNN borrows information across treatments to increase the effective sample size while retaining finite-sample error bounds and asymptotic normality. - Feasible Fusion: Constrained Joint Estimation under Structural Non-Overlap
TL;DR: Weighted fusion of observational and randomized data can violate randomized identifying restrictions under structural non-overlap; this paper jointly learns representations and predictors under orthogonal RCT moment constraints to preserve those restrictions. - Budgeted Active Experimentation for Treatment Effect Estimation from Observational and Randomized Data
TL;DR: Observational estimates can fail under policy-induced imbalance and limited overlap, while randomized trials are costly; this paper uses observational priors to allocate scarce randomization, with finite-sample, asymptotic, and minimax guarantees. - Partial Identification of Policy Values under Network Interference
TL;DR: Standard offline policy evaluation loses point identification under network interference and limited logging-policy overlap; this paper combines higher-order exposure mappings and smoothness assumptions to compute valid policy-value bounds by linear programming. - Partial Identification under High-Dimensional Potential Outcomes and Confounders via Optimal Transport
TL;DR: Optimal-transport causal bounds suffer in high dimensions, while projection remedies discard residual information; this paper combines low-dimensional signal transport with sliced-Wasserstein residual recovery to obtain tractable, more informative bounds. - MLUBench: A Benchmark for Lifelong Unlearning Evaluation in MLLMs
TL;DR: Existing unlearning benchmarks do not capture sequential multimodal forgetting; MLUBench covers 127 entities, reveals cumulative degradation and alignment failures, and introduces LUMoE to mitigate the resulting degradation. - Treatment Responder Classification with Abstention
(Spotlight, <2%).TL;DR: Causal decision methods generally lack abstention, while reject-option classifiers lack causal guarantees; TRECA links causal misclassification risk to CVaR for doubly robust responder classification and adds partial-identification and sensitivity analyses. - SpaCellAgent: A Self-evolving LLM-Based Multi-Agent Framework for Trajectory Analysis
TL;DR: Existing trajectory-analysis pipelines often require manual intervention and expertise across heterogeneous tools; SpaCellAgent automates end-to-end spatial and single-cell analysis through multi-agent planning, adaptive tool orchestration, and feedback-driven self-evolution. - Online Experimental Design With Estimation–Regret Trade-off Under Network Interference
Invited talk at POMSHK2026.TL;DR: Offline randomized designs can control estimation error but incur persistent regret in sequential network experiments; this work uses exposure mappings to unify causal estimation with reward learning and proves a Pareto-optimal estimation-regret frontier. - Design-Based Bandits Under Network Interference: Trade-Off Between Regret and Statistical Inference
Invited talk at POMSHK2026.TL;DR: The earlier MAB-N framework assumed stochastic rewards and did not provide continual inference; this work extends its estimation-regret frontier to the adversarial, design-based setting and adds anytime-valid inference through EXP3-N-CS. - Unveiling Environmental Sensitivity of Individual Gains in Influence Maximization
TL;DR: Influence-maximization models usually treat node gains as fixed; CauIM models their counterfactual dependence on neighboring activations and supplies accelerated algorithms, a generalized spread lower bound, and robustness analysis. - Active Treatment Effect Estimation via Limited Samples
TL;DR: Existing active treatment-effect estimators often lack finite-sample guarantees or optimal dependence on dimension under tight budgets; RWAS achieves near-optimal sample complexity with nonasymptotic error control and extends to network interference. - Tight Partial Identification of Causal Effects with Marginal Distribution of Unmeasured Confounders
(Spotlight, <2%).TL;DR: Proxy methods require auxiliary variables, while prior marginal-information bounds cover only extreme confounder regimes; this work derives tight closed-form bounds for any discrete confounder marginal and characterizes exactly when that information tightens identification. - Partial Identification with Proxy of Latent Confoundings via Sum-of-ratios Fractional Programming
TL;DR: Proximal causal methods rely on hard-to-verify invertibility or completeness; PI-SFP instead uses partial knowledge of the proxy-confounder transition and computes valid causal bounds through globally convergent sum-of-ratios optimization. - Robust Causal Inference for Recommender System to Overcome Noisy Confounders
TL;DR: IPS and doubly robust recommenders assume confounders are observed without error; AT-IPS models a feasible region of confounder noise and trains propensity and prediction models against the worst case, characterizing the accuracy-robustness trade-off.
Preprints & Working Papers
- Rubrics on Trial: Evolving Rubrics from a Single Query via Synthetic Pairwise Evidence
TL;DR: Existing rubric generators require external supervision or can accept plausible but untested criteria; Rubrics on Trial evolves and validates query-specific rubrics using only synthetic response pairs, without annotations or model training. - Denoised Conformal Alignment for Reliable Selection of Conditional Average Treatment Effect Predictions
TL;DR: Marginal conformal guarantees may fail on model-selected subsets, and naive CATE-error proxies lose power under heteroskedasticity; Denoised Conformal Alignment subtracts estimated conditional variance before conformal-BH calibration, yielding asymptotic FDR control with greater power. - Group Permutation Testing for Linear Models: Sharp Validity, Power Improvement, and Extension Beyond Exchangeability
TL;DR: Prior work left PALMRT's factor-two Type I bound, design-adaptive power, and validity beyond exchangeability unresolved; this group-permutation framework proves sharpness, selects design-informed groups, and derives finite-sample robustness bounds. - Orthogonal Uplift Learning with Permutation-Invariant Representations for Combinatorial Treatments
TL;DR: Categorical policy encodings ignore equivalent treatment combinations and destabilize long-tail estimates; POUL uses a permutation-invariant policy-mixture representation and an orthogonalized uplift objective to improve stability and generalization for rare or perturbed policies. - Individualized Causal Effects under Network Interference with Combinatorial Treatments
TL;DR: Prior research treats network interference, heterogeneity, and multidimensional treatments separately, creating an exponentially large joint counterfactual space; this framework combines rooted network configurations, doubly robust orthogonalization, and sparse spectral learning to estimate individualized effects. - Bounding the Largest Intersection Number of Regular Simplicial Partitions
TL;DR: Although regularity prevents degenerate simplices, the maximum number meeting at one point remained unknown; this note gives an explicit upper bound in terms of dimension and the regularity constant. - Dynamic Window-level Granger Causality of Multi-channel Time Series
TL;DR: Traditional Granger causality assumes channel relationships remain constant over time; DWGC detects dynamic cross-channel causality by testing windowed forecast errors and reweighting channels to suppress autocorrelation.
PhD Thesis
- Causal Inference under Real-world Constraints: Privacy, Robustness, and Pareto BalanceTL;DR: Idealized causal methods can fail under privacy, robustness, and competing-objective constraints; this thesis develops a framework for such real-world constraints and uses Pareto analysis to characterize attainable trade-offs.
