Aneesh Muppidi

I'm a first-year Computer Science PhD student at Stanford, advised by Chelsea Finn. I am a Rhodes Scholar, and I previously completed my undergraduate degree at Harvard.

Email  /  Scholar  /  GitHub  /  X

Aneesh Muppidi
Stanford University University of Oxford Harvard University MIT CSAIL Kempner Institute

Recent News

Sep 2026 Started my PhD in Computer Science at Stanford, working on robot learning with Chelsea Finn in the IRIS Lab.
Sep 2026 Our paper on real-time reinforcement learning, Finding the Time to Think, was accepted to NeurIPS 2026. project page
Sep 2026 Submitted LongContextVLA: Scaling Robot Memory to Hundreds of Images to ICLR 2027, work with Ji Woong Kim, Ke Wang, Ajay Sridhar, Marcel Torné, and Chelsea Finn. paper
Sep 2026 Two submissions to ICRA 2027: Mulligan, performance-guided data collection for on-robot learning (now on arXiv), and Learning Fast Dynamic Imitation Online. Mulligan paper / Fast Dynamic Imitation paper
Sep 2026 Selected as a Thrive Fellow in Stanford's Vice Provost for Graduate Education Thrive Program, a cohort program for first-year PhD students.
Sep 2026 Graduated from the University of Oxford with an MSc by Research in Engineering Science.
Jun 2026 Presented Fast TRAC at Config, Seoul.
Jul 2025 Predictive Scheduling for LLM Reasoning accepted to the ICML 2025 ES-FoMo workshop, Vancouver. project page
May 2025 Graduated from Harvard (A.B. Computer Science and Neuroscience; S.M. Computer Science) and received the Elliot and Mary Perkins Prize, awarded to one student a year for contribution to the Lowell House community.
Apr 2025 Named one of the Harvard Technology Review's Top Ten Seniors in Innovation.
Nov 2024 Named a United States Rhodes Scholar. Harvard Gazette / Forbes
Sep 2024 Fast TRAC accepted to NeurIPS 2024, after spotlight presentations at the RLC and RSS 2024 workshops. website / code
Aug 2024 Won Best Undergraduate Poster at the Kempner Institute Summer Symposium for my senior thesis work on agent discovery.
2024 Selected for the Kempner Institute's KURE and KRANIUM undergraduate research fellowships.
May 2024 Resampling-free Particle Filters in High Dimensions presented at ICRA 2024, Yokohama.

Research

Highlighted entries are ones I led or co-led. * denotes equal contribution. Google Scholar

Long-context VLA demos
LongContextVLA: Scaling Robot Memory to Hundreds of Images
Ji Woong Kim, Ke Wang, Ajay Sridhar, Aneesh Muppidi, Marcel Torné, Chelsea Finn
ICLR 2027, under review
paper

Give a robot policy more than 200 frames of visual memory, and still run it in real time.

Mulligan real-robot rollout
Mulligan: Performance-Guided Data Collection for Efficient On-Robot Learning
Lars Ankile, Perry Dong, Rohan Bhowmik, Aneesh Muppidi, David D. Yuan, Shuran Song, Chelsea Finn
arXiv 2026, under review at ICRA 2027
project page / paper

Collect robot data where the policy fails, not where it already succeeds.

Fast dynamic imitation: whiteboard scrubbing
Learning Fast Dynamic Imitation Online
John Hua Yao, Ajay Sridhar, Aneesh Muppidi, Jaden Clark, Lars Ankile, Zipeng Fu, Chelsea Finn
ICRA 2027, under review
paper

Turn position-only demonstrations into fast, forceful robot actions with twenty minutes of real-world RL.

Real-time RL: Pac-Man with an adaptive planning gate
Finding the Time to Think in Real-Time Reinforcement Learning
Aneesh Muppidi*, Firas Darwish*, Dylan Cope, João F. Henriques, Jakob N. Foerster
NeurIPS 2026
project page / code / paper

Run RL in real time, and choose how long to think. Planning agents like AlphaZero and MuZero can trade search depth for latency, and we show when they should.

Predictive scheduling
Predictive Scheduling for Efficient Inference-Time Reasoning in Large Language Models
Katrina Brown*, Aneesh Muppidi*, Rana Shahout
ICML 2025 Workshop on Efficient Systems for Foundation Models (ES-FoMo III)
project page / code / paper

Train lightweight predictors to estimate how much reasoning a prompt needs before generation, then allocate tokens where they matter most.

Fast TRAC
A Parameter-Free Optimizer for Lifelong Reinforcement Learning
Aneesh Muppidi, Zhiyu Zhang, Heng Yang
NeurIPS 2024; spotlight at the RLC 2024 and RSS 2024 workshops
project page / code / arXiv

Mitigate plasticity loss, accelerate forward transfer, and avoid policy collapse with one line of code.

Resampling-free particle filters
Resampling-Free Particle Filters in High Dimensions
Akhilan Boopathy, Aneesh Muppidi, Peggy Yang, Abhiram Iyer, William Yue, Ila Fiete
IEEE International Conference on Robotics and Automation (ICRA) 2024
arXiv

Particle filters that scale to high-dimensional state spaces without resampling, applied to object and pose estimation.

Variational agent discovery: slot attention over a Heider-Simmel scene
Variational Agent Discovery
Aneesh Muppidi, Wilka Carvalho, Samuel Gershman
Harvard senior thesis, 2025; Best Undergraduate Poster, Kempner Institute
project page / thesis

How can we discover agents using only vision?

Research Experience

IRIS Lab, Stanford University
Advisor: Chelsea Finn
PhD student, September 2026 – present
FLAIR and VGG, University of Oxford
Advisors: Jakob Foerster and João Henriques
Rhodes Scholar, MSc by Research, 2025 – 2026
Computational Robotics Lab, Harvard
Advisor: Heng Yang
Undergraduate researcher, 2023 – 2025
Computational Cognitive Neuroscience Lab, Kempner Institute at Harvard
Advisor: Samuel Gershman
KURE and KRANIUM Fellow, 2023 – 2025
Fiete Lab, MIT
Advisor: Ila Fiete
Undergraduate researcher, 2023 – 2025

Course Projects

Let's Learn Agency: Learning Emergent Agent and Non-Agent Trajectory Representations
Let's Learn Agency: Learning Emergent Agent and Non-Agent Trajectory Representations
MIT 6.8200 with Pulkit Agrawal
report
Generating Suboptimal Expert Demonstrations with Large Language Models
Generating Suboptimal Expert Demonstrations with Large Language Models
MIT 6.4212 with Russ Tedrake, advised by Lirui Wang
report / video
Rapid Learning Mechanisms and Neural Representations in Reinforcement Learning
Rapid Learning Mechanisms and Neural Representations in Reinforcement Learning
Harvard PSY 2350R with Sam Gershman, advised by Jay Henning
report / code / research notebook
Diffusion Policy for Classical Control Problems
Diffusion Policy for Classical Control Problems
Harvard ES158 with Heng Yang
report / code / slides
Visualizing Collaborative Multi-Agent Reinforcement Learning
Visualizing Collaborative Multi-Agent Reinforcement Learning
Harvard CS271 with Johanna Beyer and Hanspeter Pfister
report / slides

More

Open Source

Policy Iteration via Search Distillation: policy iteration with Gumbel MuZero, distilled into a fast neural policy in JAX/XLA; beats PPO.

NNX-Control: high-performance JAX control environments with end-to-end PPO training in a single file.

TRAC: the parameter-free optimizer from our NeurIPS 2024 paper, as a drop-in PyTorch optimizer.

Awards

United States Rhodes Scholarship, 2024
Elliot and Mary Perkins Prize, Lowell House, Harvard, 2025
Top Ten Seniors in Innovation, Harvard Technology Review, 2025
Best Undergraduate Poster, Kempner Institute Summer Symposium, 2024
KURE and KRANIUM Research Fellowships, Kempner Institute, 2024
Spotlight, RLC 2024 and RSS 2024 workshops
Thrive Fellow, Stanford Vice Provost for Graduate Education, 2026
Pechet Prize Fellowship, Harvard, 2023
John Harvard Scholarship, Harvard
Coca-Cola Scholar, 2021

Modified from Jon Barron.