Zilei Zoe Shao

Portrait of Zilei Zoe Shao
Portrait
Zilei Zoe Shao
Ph.D. Student in Computer Science
University of California, Los Angeles

About Me

I am a Ph.D. student in Computer Science at the University of California, Los Angeles, advised by Prof. Guy Van den Broeck in the StarAI Lab. I go by Zoe.

My research integrates tractable probabilistic models into modern neural architectures to improve the training efficiency, inference quality, and controllability of deep generative models. I also develop optimization techniques for probabilistic and generative models, spanning parameter learning, low-variance gradient estimation, and differentiable routing in large-scale architectures. In a related line of work, I study tokenization in Large Language Models (LLMs) as a neurosymbolic problem, showing that non-canonical tokenizations of the same input can substantially change model behavior, with the broader goal of making these large-scale systems more tractable and reliable.

Before my doctoral studies, I earned my Bachelor of Science degree in a joint major of Computer Science and Mathematics at Harvey Mudd College, where I was deeply involved in bounding-box inference within the context of model-based Reinforcement Learning under the supervision of Prof. Erin Talvitie.

Curriculum Vitae

Education

  • University of California, Los Angeles
    Ph.D. in Computer Science
    Advisor: Prof. Guy Van den Broeck
    Sep. 2024 - present
  • Harvey Mudd College
    B.S. in Joint Major of Computer Science and Mathematics
    Graduated with High Distinction, Honors in Computer Science
    GPA: 3.94 / 4.0
    Aug. 2020 - May 2024

Experience

  • Waymo LLC
    Software Engineer Intern, Predictive Planner
    Jun. 2025 - Sep. 2025
  • StarAI Lab, UCLA
    Graduate Researcher
    May 2024 - present
  • L.A.C.E. Lab, Harvey Mudd College
    Undergraduate Researcher
    Jun. 2022 - Jul. 2023

Honors & Awards

  • Oral Presentation, AISTATS 2026
    2026
  • Best Poster Award, NeurIPS 2025 SPIGM Workshop
    2025

News

2026
I served as Lead Instructor for Introduction to Computer Science at the UCLA Computer Science Summer Institute, a three-week introductory Python course for high school students.
Jul 01
I gave the talk "Token Dependencies as Vulnerability and Opportunity in Language Models" at the Harvey Mudd College Summer CS Talk.
Jun 01
Our paper Rethinking Probabilistic Circuit Parameter Learning was accepted to AISTATS 2026 as an Oral.
Jan 01
2025
Zero-Variance Gradients for Variational Autoencoders received a Best Poster Award at the NeurIPS 2025 SPIGM Workshop.
Dec 06
I joined Waymo as a Software Engineer Intern on the Predictive Planner team (Jun–Sep 2025).
Jun 01
Our paper Adversarial Tokenization was accepted to ACL 2025 (Main Conference).
May 01
I gave a talk on Adversarial Tokenization at the Shang Data Lab, UCSD.
Apr 14
2024
I started my Ph.D. in Computer Science at UCLA.
Sep 01
Bounding-Box Inference for Error-Aware Model-Based Reinforcement Learning was published in the Reinforcement Learning Journal (RLC 2024).
Aug 09

Selected Publications (view all )

Breaking the Factorization Barrier in Diffusion Language Models

Breaking the Factorization Barrier in Diffusion Language Models

Ian Li*, Zilei Shao*, Benjie Wang, Rose Yu, Guy Van den Broeck, Anji Liu (* equal contribution)

International Conference on Machine Learning (ICML) 2026

Diffusion language models generate many tokens at once but assume they are independent, which costs coherence. CoDD replaces that factorized output with a lightweight tractable inference layer, keeping quality even at very few decoding steps.

Breaking the Factorization Barrier in Diffusion Language Models

Ian Li*, Zilei Shao*, Benjie Wang, Rose Yu, Guy Van den Broeck, Anji Liu (* equal contribution)

International Conference on Machine Learning (ICML) 2026

Diffusion language models generate many tokens at once but assume they are independent, which costs coherence. CoDD replaces that factorized output with a lightweight tractable inference layer, keeping quality even at very few decoding steps.

Rethinking Probabilistic Circuit Parameter Learning

Rethinking Probabilistic Circuit Parameter Learning

Anji Liu, Zilei Shao, Guy Van den Broeck

International Conference on Artificial Intelligence and Statistics (AISTATS) 2026 Oral

Full-batch EM for probabilistic circuits is too slow on large datasets, and mini-batch variants overfit each batch. anemone is a mini-batch EM algorithm with an implicit per-parameter adaptive learning rate that converges faster and reaches better likelihoods.

Rethinking Probabilistic Circuit Parameter Learning

Anji Liu, Zilei Shao, Guy Van den Broeck

International Conference on Artificial Intelligence and Statistics (AISTATS) 2026 Oral

Full-batch EM for probabilistic circuits is too slow on large datasets, and mini-batch variants overfit each batch. anemone is a mini-batch EM algorithm with an implicit per-parameter adaptive learning rate that converges faster and reaches better likelihoods.

Zero-Variance Gradients for Variational Autoencoders

Zero-Variance Gradients for Variational Autoencoders

Zilei Shao, Anji Liu, Guy Van den Broeck

NeurIPS Workshop on Structured Probabilistic Inference & Generative Modeling (SPIGM) 2025 Best Poster Award

Training a VAE normally requires noisy Monte Carlo gradient estimates. This work shows that with a suitably restricted decoder the expected ELBO can be computed analytically, giving gradients with zero estimation variance and better early optimization.

Zero-Variance Gradients for Variational Autoencoders

Zilei Shao, Anji Liu, Guy Van den Broeck

NeurIPS Workshop on Structured Probabilistic Inference & Generative Modeling (SPIGM) 2025 Best Poster Award

Training a VAE normally requires noisy Monte Carlo gradient estimates. This work shows that with a suitably restricted decoder the expected ELBO can be computed analytically, giving gradients with zero estimation variance and better early optimization.

Adversarial Tokenization

Adversarial Tokenization

Renato L. Geh*, Zilei Shao*, Guy Van den Broeck (* equal contribution)

Annual Meeting of the Association for Computational Linguistics (ACL), Main Conference, 2025

The same string can be tokenized in exponentially many valid ways, and LLMs still understand the non-canonical ones. Simply re-tokenizing a prompt, without changing a single character, is enough to bypass safety alignment in state-of-the-art models.

Adversarial Tokenization

Renato L. Geh*, Zilei Shao*, Guy Van den Broeck (* equal contribution)

Annual Meeting of the Association for Computational Linguistics (ACL), Main Conference, 2025

The same string can be tokenized in exponentially many valid ways, and LLMs still understand the non-canonical ones. Simply re-tokenizing a prompt, without changing a single character, is enough to bypass safety alignment in state-of-the-art models.