About
My research interests include generative modeling and multi-agent LLM systems for scientific data curation. I have worked on discrete diffusion models that unify diffusion and autoregressive generation (CaDDi, NeurIPS 2025) and that learn distributions over finite symmetric groups (Soft-Rank Diffusion, ICML 2026). I also build LLM-based multi-agent systems for automating scientific data curation at scale.
- – present
Ph.D. in Computer Science, Yale University · advised by Dr. David van Dijk
B.S. in Honors Mathematics, minor in Computer Science, University of Michigan, Ann Arbor
News
newest first · 8 entries
2 papers accepted to NeurIPS 2026.
I joined Chan Zuckerberg Biohub, Inc as an AI Research Scientist Intern for Summer 2026.
1 paper accepted to the Journal of Computational Physics.
2 papers accepted to ICML 2026.
4 earlier
1 paper accepted to NeurIPS 2025.
1 paper accepted to AI4MATH@ICML 2025.
I received the Fan Family Fellowship of Yale University.
1 paper accepted to ICLR 2025.
Selected
4 of 11 · * equal contributionThumbnails are schematic animations of each model; tokens and values are illustrative, not results. Hover or tap to replay.
-
ICML 2026Soft-Rank Diffusion
Learning Permutation Distributions via Reflected Diffusion on Ranks
Sizhuang He*, Yangtian Zhang*, Shiyang Zhang and David van Dijk
A discrete diffusion framework for distributions over permutations on Sn: discrete ranks are relaxed into soft ranks for a smoother forward process, and contextualized generalized Plackett–Luce (cGPL) denoisers reverse it, with the largest gains on long sequences.
@misc{he2026learningpermutationdistributionsreflected, title={Learning Permutation Distributions via Reflected Diffusion on Ranks}, author={Sizhuang He and Yangtian Zhang and Shiyang Zhang and David van Dijk}, year={2026}, eprint={2603.17353}, archivePrefix={arXiv}, primaryClass={cs.LG}, url={https://arxiv.org/abs/2603.17353},}
The finite symmetric group Sn provides a natural domain for permutations, yet learning probability distributions on Sn is challenging due to its factorially growing size and discrete, non-Euclidean structure. Recent permutation diffusion methods define forward noising via shuffle-based random walks (e.g., riffle shuffles) and learn reverse transitions with Plackett–Luce (PL) variants, but the resulting trajectories can be abrupt and increasingly hard to denoise as n grows. We propose Soft-Rank Diffusion, a discrete diffusion framework that replaces shuffle-based corruption with a structured soft-rank forward process: we lift permutations to a continuous latent representation of order by relaxing discrete ranks into soft ranks, yielding smoother and more tractable trajectories. For the reverse process, we introduce contextualized generalized Plackett–Luce (cGPL) denoisers that generalize prior PL-style parameterizations and improve expressivity for sequential decision structures. Experiments on sorting and combinatorial optimization benchmarks show that Soft-Rank Diffusion consistently outperforms prior diffusion baselines, with particularly strong gains in long-sequence and intrinsically sequential settings.

-
Preprint 2025Cell2Sentence
Scaling Large Language Models for Next-Generation Single-Cell Analysis
Syed Rizvi*, Daniel Levine*, Aakash Patel*, Shiyang Zhang*, Eric Wang*, Sizhuang He, David Zhang, Cerise Tang, Zhuoyang Lyu, Rayyan Darji, Marlene Li, Emily Sun, David Jeong, Lawrence Zhao, Jennifer Kwan, David Braun, Brian Hafler, Jeffrey Ishizuka, Rahul Dhodapkar, Hattie Chung, Shekoofeh Azizi, Bryan Perozzi, and David van Dijk
Scales the Cell2Sentence framework to 27B-parameter LLMs trained on over a billion tokens of transcriptomic data, biological text, and metadata, reaching state-of-the-art predictive and generative performance on multicellular analyses.
@article {Rizvi2025.04.14.648850, author = {Rizvi, Syed Asad and Levine, Daniel and Patel, Aakash and Zhang, Shiyang and Wang, Eric and He, Sizhuang and Zhang, David and Tang, Cerise and Lyu, Zhuoyang and Darji, Rayyan and Li, Marlene and Sun, Emily and Jeong, David and Zhao, Lawrence and Kwan, Jennifer and Braun, David and Hafler, Brian and Ishizuka, Jeffrey and Dhodapkar, Rahul and Chung, Hattie and Azizi, Shekoofeh and Perozzi, Bryan and van Dijk, David}, title = {Scaling Large Language Models for Next-Generation Single-Cell Analysis}, elocation-id = {2025.04.14.648850}, year = {2025}, doi = {10.1101/2025.04.14.648850}, publisher = {Cold Spring Harbor Laboratory}, abstract = {Single-cell RNA sequencing has transformed our understanding of cellular diversity, yet current single-cell foundation models (scFMs) remain limited in their scalability, flexibility across diverse tasks, and ability to natively integrate textual information. In this work, we build upon the Cell2Sentence (C2S) framework, which represents scRNA-seq profiles as textual "cell sentences," to train Large Language Models (LLMs) on a corpus comprising over one billion tokens of transcriptomic data, biological text, and metadata. By scaling model size to 27 billion parameters, we observe consistent improvements in predictive and generative capabilities, as well as the capacity for advanced downstream tasks requiring synthesis of information across multicellular contexts. Through targeted fine-tuning supported by modern reinforcement learning techniques, our approach excels in tasks such as perturbation response prediction, natural language interpretation, and complex biological reasoning. By unifying transcriptomic and textual data at unprecedented scales, this approach not only surpasses both specialized single-cell models and general-purpose LLMs, but also establishes a powerful platform for next-generation single-cell analysis, paving the way for the development of "virtual cells."Competing Interest StatementThe authors have declared no competing interest.}, URL = {https://www.biorxiv.org/content/early/2025/04/17/2025.04.14.648850}, eprint = {https://www.biorxiv.org/content/early/2025/04/17/2025.04.14.648850.full.pdf}, journal = {bioRxiv}}
Single-cell RNA sequencing has transformed our understanding of cellular diversity, yet current single-cell foundation models (scFMs) remain limited in their scalability, flexibility across diverse tasks, and ability to natively integrate textual information. In this work, we build upon the Cell2Sentence (C2S) framework, which represents scRNA-seq profiles as textual 'cell sentences,' to train Large Language Models (LLMs) on a corpus comprising over one billion tokens of transcriptomic data, biological text, and metadata. By scaling model size to 27 billion parameters, we observe consistent improvements in predictive and generative capabilities, as well as the capacity for advanced downstream tasks requiring synthesis of information across multicellular contexts. Through targeted fine-tuning supported by modern reinforcement learning techniques, our approach excels in tasks such as perturbation response prediction, natural language interpretation, and complex biological reasoning. By unifying transcriptomic and textual data at unprecedented scales, this approach not only surpasses both specialized single-cell models and general-purpose LLMs, but also establishes a powerful platform for next-generation single-cell analysis, paving the way for the development of 'virtual cells.'

-
NeurIPS 2025CaDDi
Non-Markovian Discrete Diffusion with Causal Language Models
Yangtian Zhang*, Sizhuang He*, Daniel Levine, Lawrence Zhao, David Zhang, Syed Rizvi, Emanuele Zappala, Rex Ying, and David van Dijk
Lifts the Markov constraint in discrete diffusion by conditioning on the entire generative trajectory; causal language models become a special case, so pretrained LLM weights are reused with no architectural changes.
@misc{zhang2025nonmarkoviandiscretediffusioncausal, title={Non-Markovian Discrete Diffusion with Causal Language Models}, author={Yangtian Zhang and Sizhuang He and Daniel Levine and Lawrence Zhao and David Zhang and Syed A Rizvi and Emanuele Zappala and Rex Ying and David van Dijk}, year={2025}, eprint={2502.09767}, archivePrefix={arXiv}, primaryClass={cs.LG}, url={https://arxiv.org/abs/2502.09767},}
Discrete diffusion models offer a flexible, controllable approach to structured sequence generation, yet they still lag behind causal language models in expressive power. A key limitation lies in their reliance on the Markovian assumption, which restricts each step to condition only on the current state, leading to potential uncorrectable error accumulation. In this paper, we introduce CaDDi, a discrete diffusion model that conditions on the entire generative trajectory, thereby lifting the Markov constraint and allowing the model to revisit and improve past states. By unifying sequential (causal) and temporal (diffusion) reasoning in a single non-Markovian transformer, CaDDi also treats standard causal language models as a special case and permits the direct reuse of pretrained LLM weights with no architectural changes. Empirically, CaDDi outperforms state-of-the-art discrete diffusion baselines on natural-language benchmarks, substantially narrowing the remaining gap to large autoregressive transformers.

-
ICLR 2025Edge of Chaos
Intelligence at the Edge of Chaos
Shiyang Zhang*, Aakash Patel*, Syed Rizvi, Nianchen Liu, Sizhuang He, Amin Karbasi, Emanuele Zappala, and David van Dijk
Training LLMs on elementary cellular automata of varying complexity reveals a sweet spot of data complexity that maximizes downstream reasoning and prediction.
@inproceedings{ICLR2025_d791394d, author = {Zhang, Shiyang and Patel, Aakash and Rizvi, Syed and Liu, Nianchen and He, Sizhuang and Karbasi, Amin and Zappala, Emanuele and van Dijk, David}, booktitle = {International Conference on Representation Learning}, editor = {Y. Yue and A. Garg and N. Peng and F. Sha and R. Yu}, pages = {86576--86592}, title = {Intelligence at the Edge of Chaos}, url = {https://proceedings.iclr.cc/paper_files/paper/2025/file/d791394d32c428aecc7a5b101fb47799-Paper-Conference.pdf}, volume = {2025}, year = {2025}}
We explore the emergence of intelligent behavior in artificial systems by investigating how the complexity of rule-based systems influences the capabilities of models trained to predict these rules. Our study focuses on elementary cellular automata (ECA), simple yet powerful one-dimensional systems that generate behaviors ranging from trivial to highly complex. By training distinct Large Language Models (LLMs) on different ECAs, we evaluated the relationship between the complexity of the data generated by the rules and the models' ability to learn effective general representations, as reflected in their performance on downstream tasks. Our findings reveal that models trained on more complex data exhibit greater predictive ability, as demonstrated by their performance on reasoning and chess move prediction tasks. Both uniform and periodic systems, and often also highly chaotic systems, resulted in poorer downstream performance, highlighting a sweet spot of complexity conducive to intelligence. We conjecture that intelligence arises from the ability to predict complexity and that creating intelligence may require only exposure to complexity.

Publications
11 papers
Schematic thumbnails · illustrative, not results
applied
-
01NeurIPS 2026
MoRSE: Task-Oriented Multi-Agent System with Mixture of Role-Subtask Experts
Peiwen Li, Shiyang Zhang, Yangtian Zhang, Sizhuang He, David van Dijk, and Rex Ying
@misc{li2026morsetaskorientedmultiagentsystem, title={MoRSE: Task-Oriented Multi-Agent System with Mixture of Role-Subtask Experts}, author={Peiwen Li and Shiyang Zhang and Yangtian Zhang and Sizhuang He and David van Dijk and Rex Ying}, year={2026}, eprint={2608.09251}, archivePrefix={arXiv}, primaryClass={cs.MA}, url={https://arxiv.org/abs/2608.09251},}
Large language model-based multi-agent systems have recently shown strong potential for complex, long-horizon tasks. However, existing methods mainly rely on coarse prompt-level differentiation without parameter adaptation for diverse subtasks, resulting in insufficient inter-agent heterogeneity and limited specialized capability that bottleneck performance on tasks with complex requirements. To address this, we introduce a Task-Oriented Multi-Agent System with Mixture of Role-Subtask Experts (MoRSE) that distinguishes agents with (role, subtask)-conditional specialization at both the task structure and parameter levels. To make agents' responsibility explicit at the task structure level, we formulate a task-oriented multi-agent system that decomposes each task into a dependency-aware Directed Acyclic Graph of subtasks and assigns each agent a specific (role, subtask), introducing task-level specialization across collaborating agents. Additionally, to address the diverse role and subtask parameter adaptation demands, we propose a dynamic Mixture of (role, subtask) LoRA Experts module with a prototype-based semantic router for subtasks, augmenting agents with parameter-level specialization on a shared LLM substrate cost-effectively. Then, to co-optimize experts and router stably under sparse task rewards, we further propose a hierarchical group-relative policy optimization with two-layer credit assignment that isolates expert updates from the cross-route variance introduced by routing decisions, disentangling expert quality from routing quality. Experiments on code-generation benchmarks across three backbones demonstrate the effectiveness of our approach, with improvements in both whole-task and step-wise performance, and the gains from trained specialization generalize across held-out task categories and domains.
MoRSE
-
02ICML 2026
Learning Permutation Distributions via Reflected Diffusion on Ranks
Sizhuang He*, Yangtian Zhang*, Shiyang Zhang and David van Dijk
@misc{he2026learningpermutationdistributionsreflected, title={Learning Permutation Distributions via Reflected Diffusion on Ranks}, author={Sizhuang He and Yangtian Zhang and Shiyang Zhang and David van Dijk}, year={2026}, eprint={2603.17353}, archivePrefix={arXiv}, primaryClass={cs.LG}, url={https://arxiv.org/abs/2603.17353},}
The finite symmetric group Sn provides a natural domain for permutations, yet learning probability distributions on Sn is challenging due to its factorially growing size and discrete, non-Euclidean structure. Recent permutation diffusion methods define forward noising via shuffle-based random walks (e.g., riffle shuffles) and learn reverse transitions with Plackett–Luce (PL) variants, but the resulting trajectories can be abrupt and increasingly hard to denoise as n grows. We propose Soft-Rank Diffusion, a discrete diffusion framework that replaces shuffle-based corruption with a structured soft-rank forward process: we lift permutations to a continuous latent representation of order by relaxing discrete ranks into soft ranks, yielding smoother and more tractable trajectories. For the reverse process, we introduce contextualized generalized Plackett–Luce (cGPL) denoisers that generalize prior PL-style parameterizations and improve expressivity for sequential decision structures. Experiments on sorting and combinatorial optimization benchmarks show that Soft-Rank Diffusion consistently outperforms prior diffusion baselines, with particularly strong gains in long-sequence and intrinsically sequential settings.
Soft-Rank Diffusion

-
03ICML 2026
STRIDE: Post-Training LLMs to Reason and Refine Bio-Sequences via Edit Trajectories
Daiheng Zhang, Shiyang Zhang, Sizhuang He, Yangtian Zhang and David van Dijk
@misc{zhang2026strideposttrainingllmsreason, title={STRIDE: Post-Training LLMs to Reason and Refine Bio-Sequences via Edit Trajectories}, author={Daiheng Zhang and Shiyang Zhang and Sizhuang He and Yangtian Zhang and Syed Asad Rizvi and David van Dijk}, year={2026}, eprint={2603.03573}, archivePrefix={arXiv}, primaryClass={cs.CE}, url={https://arxiv.org/abs/2603.03573},}
Discrete biological sequence optimization demands iterative refinement while satisfying strict syntactic constraints. Diffusion-based approaches provide strong progressive refinement but are not naturally aligned with discrete, grammar-constrained edit operations, whereas autoregressive LLMs readily produce valid sequences yet often lack explicit long-horizon planning. To close this gap, we introduce STRIDE (Sequence Trajectory Refinement via Internalized Denoising Emulation), a post-training framework that recasts optimization as an intrinsic reasoning problem in edit space. Rather than relying on external agentic search loops, STRIDE trains an LLM to emit a full trajectory of atomic edits as explicit Chain-of-Thought, effectively internalizing a trajectory-based refinement policy under discrete constraints. We instantiate STRIDE with a curriculum that combines supervised fine-tuning on Levenshtein-aligned shortest-edit demonstrations with GRPO-style reinforcement learning (and variants) to align edit trajectories with task rewards. Across protein and molecule optimization benchmarks, STRIDE consistently outperforms a diverse set of baselines, while producing candidates that maintain high structural validity and achieve improved target properties.
STRIDE
-
04NeurIPS 2026
FLUX: Geometry-Aware Longitudinal Flow Matching with Mixture of Experts
Josue Ortega Caro, Yongxu Zhang, Hannah M Batchelor, Sizhuang He, Jessica Cardin and Shreya Saxena
@misc{caro2026fluxgeometryawarelongitudinalflow, title={FLUX: Geometry-Aware Longitudinal Flow Matching with Mixture of Experts}, author={Josue Ortega Caro and Yongxu Zhang and Hannah M Batchelor and Sizhuang He and Jessica Cardin and Shreya Saxena}, year={2026}, eprint={2605.08648}, archivePrefix={arXiv}, primaryClass={cs.LG}, url={https://arxiv.org/abs/2605.08648},}
Many biological systems evolve through continuous local dynamics while switching between latent regimes defined by learning, stimulus context, internal state, or developmental stage. These processes are often observed only as unpaired longitudinal snapshots: the same cells, neurons, or animals are not tracked as matched trajectories, even though population states are sampled across successive stages. This creates two coupled challenges. First, trajectories must respect curved low-dimensional manifolds embedded in high-dimensional biological measurements. Second, the model must identify when the transport mechanism itself changes. We introduce FLUX (FLow matching for Unpaired longitudinal data with miXture-of-experts), a geometry-aware longitudinal flow-matching framework for joint transport modeling and unsupervised regime discovery. FLUX learns a data-dependent metric from pooled labeled and unlabeled observations, uses that metric to construct geometry-aware conditional paths between adjacent marginals, and decomposes the resulting velocity field into sparse expert vector fields selected by a Straight-Through Gumbel-Softmax router. Across manifold controls, a regime-switching Lorenz system, widefield cortical calcium imaging during associative learning, and embryoid body single-cell differentiation, FLUX reconstructs longitudinal transport while recovering interpretable regime structure. Ablations show that mixture-of-experts routing alone is insufficient: FLUX without geometric learning can fit local transport but fails or weakens regime discovery when regimes are encoded in local dynamics. These results suggest that geometry-aware velocity decomposition provides a general strategy for discovering latent biological state transitions from unpaired longitudinal snapshots.
FLUX
-
05NeurIPS 2025
Non-Markovian Discrete Diffusion with Causal Language Models
Yangtian Zhang*, Sizhuang He*, Daniel Levine, Lawrence Zhao, David Zhang, Syed Rizvi, Emanuele Zappala, Rex Ying, and David van Dijk
@misc{zhang2025nonmarkoviandiscretediffusioncausal, title={Non-Markovian Discrete Diffusion with Causal Language Models}, author={Yangtian Zhang and Sizhuang He and Daniel Levine and Lawrence Zhao and David Zhang and Syed A Rizvi and Emanuele Zappala and Rex Ying and David van Dijk}, year={2025}, eprint={2502.09767}, archivePrefix={arXiv}, primaryClass={cs.LG}, url={https://arxiv.org/abs/2502.09767},}
Discrete diffusion models offer a flexible, controllable approach to structured sequence generation, yet they still lag behind causal language models in expressive power. A key limitation lies in their reliance on the Markovian assumption, which restricts each step to condition only on the current state, leading to potential uncorrectable error accumulation. In this paper, we introduce CaDDi, a discrete diffusion model that conditions on the entire generative trajectory, thereby lifting the Markov constraint and allowing the model to revisit and improve past states. By unifying sequential (causal) and temporal (diffusion) reasoning in a single non-Markovian transformer, CaDDi also treats standard causal language models as a special case and permits the direct reuse of pretrained LLM weights with no architectural changes. Empirically, CaDDi outperforms state-of-the-art discrete diffusion baselines on natural-language benchmarks, substantially narrowing the remaining gap to large autoregressive transformers.
CaDDi

-
06JCP 2026
TANTE: Time-Adaptive Operator Learning via Neural Taylor Expansion
Zhikai Wu, Sifan Wang, Shiyang Zhang, Sizhuang He, Min Zhu, Anran Jiao, Lu Lu, and David van Dijk
@misc{wu2025tantetimeadaptiveoperatorlearning, title={TANTE: Time-Adaptive Operator Learning via Neural Taylor Expansion}, author={Zhikai Wu and Sifan Wang and Shiyang Zhang and Sizhuang He and Min Zhu and Anran Jiao and Lu Lu and David van Dijk}, year={2025}, eprint={2502.08574}, archivePrefix={arXiv}, primaryClass={cs.LG}, url={https://arxiv.org/abs/2502.08574},}
Operator learning for time-dependent partial differential equations (PDEs) has seen rapid progress in recent years, enabling efficient approximation of complex spatiotemporal dynamics. However, most existing methods rely on fixed time step sizes during rollout, which limits their ability to adapt to varying temporal complexity and often leads to error accumulation. Here, we propose the Time-Adaptive Transformer with Neural Taylor Expansion (TANTE), a novel operator-learning framework that produces continuous-time predictions with adaptive step sizes. TANTE predicts future states by performing a Taylor expansion at the current state, where neural networks learn both the higher-order temporal derivatives and the local radius of convergence. This allows the model to dynamically adjust its rollout based on the local behavior of the solution, thereby reducing cumulative error and improving computational efficiency. We demonstrate the effectiveness of TANTE across a wide range of PDE benchmarks, achieving superior accuracy and adaptability compared to fixed-step baselines, delivering accuracy gains of 60–80% and speed-ups of 30–40% at inference time. The code is publicly available at https://github.com/zwu88/TANTE for transparency and reproducibility.
TANTE

-
07AI4MATH@ICML 2025
COAST: Intelligent Time-Adaptive Neural Operators
Zhikai Wu, Shiyang Zhang, Sizhuang He, Sifan Wang, Min Zhu, Anran Jiao, Lu Lu, and David van Dijk
@misc{wu2025coastintelligenttimeadaptiveneural, title={COAST: Intelligent Time-Adaptive Neural Operators}, author={Zhikai Wu and Shiyang Zhang and Sizhuang He and Sifan Wang and Min Zhu and Anran Jiao and Lu Lu and David van Dijk}, year={2025}, eprint={2502.08574}, archivePrefix={arXiv}, primaryClass={cs.LG}, url={https://arxiv.org/abs/2502.08574},}
We introduce Causal Operator with Adaptive Solver Transformer (COAST), a novel neural operator learning method that leverages a causal language model (CLM) framework to dynamically adapt time steps. Our method predicts both the evolution of a system and its optimal time step, intelligently balancing computational efficiency and accuracy. We find that COAST generates variable step sizes that correlate with the underlying system intrinsicities, both within and across dynamical systems. Within a single trajectory, smaller steps are taken in regions of high complexity, while larger steps are employed in simpler regions. Across different systems, more complex dynamics receive more granular time steps. Benchmarked on diverse systems with varied dynamics, COAST consistently outperforms state-of-the-art methods, achieving superior performance in both efficiency and accuracy. This work underscores the potential of CLM-based intelligent adaptive solvers for scalable operator learning of dynamical systems.
COAST

-
08Preprint 2025
Scaling Large Language Models for Next-Generation Single-Cell Analysis
Syed Rizvi*, Daniel Levine*, Aakash Patel*, Shiyang Zhang*, Eric Wang*, Sizhuang He, David Zhang, Cerise Tang, Zhuoyang Lyu, Rayyan Darji, Marlene Li, Emily Sun, David Jeong, Lawrence Zhao, Jennifer Kwan, David Braun, Brian Hafler, Jeffrey Ishizuka, Rahul Dhodapkar, Hattie Chung, Shekoofeh Azizi, Bryan Perozzi, and David van Dijk
@article {Rizvi2025.04.14.648850, author = {Rizvi, Syed Asad and Levine, Daniel and Patel, Aakash and Zhang, Shiyang and Wang, Eric and He, Sizhuang and Zhang, David and Tang, Cerise and Lyu, Zhuoyang and Darji, Rayyan and Li, Marlene and Sun, Emily and Jeong, David and Zhao, Lawrence and Kwan, Jennifer and Braun, David and Hafler, Brian and Ishizuka, Jeffrey and Dhodapkar, Rahul and Chung, Hattie and Azizi, Shekoofeh and Perozzi, Bryan and van Dijk, David}, title = {Scaling Large Language Models for Next-Generation Single-Cell Analysis}, elocation-id = {2025.04.14.648850}, year = {2025}, doi = {10.1101/2025.04.14.648850}, publisher = {Cold Spring Harbor Laboratory}, abstract = {Single-cell RNA sequencing has transformed our understanding of cellular diversity, yet current single-cell foundation models (scFMs) remain limited in their scalability, flexibility across diverse tasks, and ability to natively integrate textual information. In this work, we build upon the Cell2Sentence (C2S) framework, which represents scRNA-seq profiles as textual "cell sentences," to train Large Language Models (LLMs) on a corpus comprising over one billion tokens of transcriptomic data, biological text, and metadata. By scaling model size to 27 billion parameters, we observe consistent improvements in predictive and generative capabilities, as well as the capacity for advanced downstream tasks requiring synthesis of information across multicellular contexts. Through targeted fine-tuning supported by modern reinforcement learning techniques, our approach excels in tasks such as perturbation response prediction, natural language interpretation, and complex biological reasoning. By unifying transcriptomic and textual data at unprecedented scales, this approach not only surpasses both specialized single-cell models and general-purpose LLMs, but also establishes a powerful platform for next-generation single-cell analysis, paving the way for the development of "virtual cells."Competing Interest StatementThe authors have declared no competing interest.}, URL = {https://www.biorxiv.org/content/early/2025/04/17/2025.04.14.648850}, eprint = {https://www.biorxiv.org/content/early/2025/04/17/2025.04.14.648850.full.pdf}, journal = {bioRxiv}}
Single-cell RNA sequencing has transformed our understanding of cellular diversity, yet current single-cell foundation models (scFMs) remain limited in their scalability, flexibility across diverse tasks, and ability to natively integrate textual information. In this work, we build upon the Cell2Sentence (C2S) framework, which represents scRNA-seq profiles as textual 'cell sentences,' to train Large Language Models (LLMs) on a corpus comprising over one billion tokens of transcriptomic data, biological text, and metadata. By scaling model size to 27 billion parameters, we observe consistent improvements in predictive and generative capabilities, as well as the capacity for advanced downstream tasks requiring synthesis of information across multicellular contexts. Through targeted fine-tuning supported by modern reinforcement learning techniques, our approach excels in tasks such as perturbation response prediction, natural language interpretation, and complex biological reasoning. By unifying transcriptomic and textual data at unprecedented scales, this approach not only surpasses both specialized single-cell models and general-purpose LLMs, but also establishes a powerful platform for next-generation single-cell analysis, paving the way for the development of 'virtual cells.'
Cell2Sentence

-
09ICLR 2025
Intelligence at the Edge of Chaos
Shiyang Zhang*, Aakash Patel*, Syed Rizvi, Nianchen Liu, Sizhuang He, Amin Karbasi, Emanuele Zappala, and David van Dijk
@inproceedings{ICLR2025_d791394d, author = {Zhang, Shiyang and Patel, Aakash and Rizvi, Syed and Liu, Nianchen and He, Sizhuang and Karbasi, Amin and Zappala, Emanuele and van Dijk, David}, booktitle = {International Conference on Representation Learning}, editor = {Y. Yue and A. Garg and N. Peng and F. Sha and R. Yu}, pages = {86576--86592}, title = {Intelligence at the Edge of Chaos}, url = {https://proceedings.iclr.cc/paper_files/paper/2025/file/d791394d32c428aecc7a5b101fb47799-Paper-Conference.pdf}, volume = {2025}, year = {2025}}
We explore the emergence of intelligent behavior in artificial systems by investigating how the complexity of rule-based systems influences the capabilities of models trained to predict these rules. Our study focuses on elementary cellular automata (ECA), simple yet powerful one-dimensional systems that generate behaviors ranging from trivial to highly complex. By training distinct Large Language Models (LLMs) on different ECAs, we evaluated the relationship between the complexity of the data generated by the rules and the models' ability to learn effective general representations, as reflected in their performance on downstream tasks. Our findings reveal that models trained on more complex data exhibit greater predictive ability, as demonstrated by their performance on reasoning and chess move prediction tasks. Both uniform and periodic systems, and often also highly chaotic systems, resulted in poorer downstream performance, highlighting a sweet spot of complexity conducive to intelligence. We conjecture that intelligence arises from the ability to predict complexity and that creating intelligence may require only exposure to complexity.
Edge of Chaos

-
10Preprint 2024
CaLMFlow: Flow Matching using Causal Language Models
Sizhuang He*, Daniel Levine*, Ivan Vrkic, Marco Bressana, David Zhang, Syed Rizvi, Yangtian Zhang, Emanuele Zappala, and David van Dijk
@misc{he2024calmflowvolterraflowmatching, title={CaLMFlow: Volterra Flow Matching using Causal Language Models}, author={Sizhuang He and Daniel Levine and Ivan Vrkic and Marco Francesco Bressana and David Zhang and Syed Asad Rizvi and Yangtian Zhang and Emanuele Zappala and David van Dijk}, year={2024}, eprint={2410.05292}, archivePrefix={arXiv}, primaryClass={cs.LG}, url={https://arxiv.org/abs/2410.05292},}
We introduce CaLMFlow, a novel framework that recasts flow matching as a Volterra integral equation (VIE) while leveraging causal language models (CLMs) for continuous data generation. Although integral equations have previously been studied for nonlocal operator learning, we are the first to apply them explicitly to flow matching. By discretizing both time and space, CaLMFlow transforms continuous trajectories into a sequence-based representation, allowing standard large language model architectures to capture long-range dependencies and flexibly incorporate textual prompts. In experiments on synthetic benchmarks and a single-cell perturbation response task, CaLMFlow demonstrates improved stability and scalability over ODE-based methods in high-dimensional regimes. These findings suggest that combining an integral-equation formulation with CLMs provides a promising, context-aware paradigm for generative modeling.
CaLMFlow

-
11Preprint 2023
Operator Learning Meets Numerical Analysis: Improving Neural Networks through Iterative Methods
Emanuele Zappala, Daniel Levine, Sizhuang He, Syed Rizvi, Sacha Lévy, and David van Dijk
@misc{zappala2023operatorlearningmeetsnumerical, title={Operator Learning Meets Numerical Analysis: Improving Neural Networks through Iterative Methods}, author={Emanuele Zappala and Daniel Levine and Sizhuang He and Syed Rizvi and Sacha Levy and David van Dijk}, year={2023}, eprint={2310.01618}, archivePrefix={arXiv}, primaryClass={cs.LG}, url={https://arxiv.org/abs/2310.01618},}
Deep neural networks, despite their success in numerous applications, often function without established theoretical foundations. In this paper, we bridge this gap by drawing parallels between deep learning and classical numerical analysis. By framing neural networks as operators with fixed points representing desired solutions, we develop a theoretical framework grounded in iterative methods for operator equations. Under defined conditions, we present convergence proofs based on fixed point theory. We demonstrate that popular architectures, such as diffusion models and AlphaFold, inherently employ iterative operator learning. Empirical assessments highlight that performing iterations through network operators improves performance. We also introduce an iterative graph neural network, PIGN, that further demonstrates benefits of iterations. Our work aims to enhance the understanding of deep learning by merging insights from numerical analysis, potentially guiding the design of future networks with clearer theoretical underpinnings and improved performance.
Iterative Methods

Service
Journal Reviewer
- ACM Computing Surveys
- Transactions on Machine Learning Research
Conference Reviewer
- AAAI Conference on Artificial Intelligence2027
- Conference on Neural Information Processing Systems2026
- International Conference on Machine Learning2026
- International Conference on Learning Representations2026
- AI4MATH Workshop at the International Conference on Machine Learning2025