Anmol Kabra

anmol (at) cs.cornell.edu   •   CV   •   twitter   •   google scholar   •   linkedin   •   bsky   •   github

images/site/anmolkabra.jpg

I am a Computer Science PhD student at Cornell University, working on LLM post-training and AI for Science. I am fortunate to be advised by Kilian Weinberger and affiliated with AI Materials Institute.

Scientific discovery requires AI agents to generalize beyond their training distribution, which demands two things: transferable reasoning skills and specialized domain knowledge. Because curating training data for either is expensive and hard to scale, I train LLMs on synthetic data and RL environments that teach transferable skills such as knowledge composition, question decomposition, and retrieval. Along the second axis of specializing agents, I study fine-tuning and context engineering to ground LLMs in scientific domains underrepresented in pretraining.

Recently as an intern at Snorkel AI, I worked on Terminal-Bench-Science and synthetic RL training environments for scientific AI agents. Previously, I was an AI/ML Quant Intern at Bloomberg’s AI Engineering team, where I prototyped tool-use agents for the ASKB<GO> function. I was also a Research Engineer at ASAPP, working on LLMs, privacy, and anomaly detection with Ethan Elenberg and Kilian Weinberger.

I received my BS in Computer Science from Cornell University and MS from Toyota Technological Institute at Chicago (TTIC). At Cornell, I was named a Merrill Presidential Scholar for my undergraduate research with Carla Gomes and Kilian Weinberger. I was generously supported by the Tata Scholarship and Telluride Scholarship.

Fun fact: I juggle more hobbies than I can juggle number of balls 🤹‍♂️


news


2026 Oct
2026 May
2026 Mar
2026 Jan
2025
2024

selected papers


* equal contributions

  1. PhantomEnvironments: Training LLM Agents in Fictional Worlds
    Under Review (2026). Oral presentation at COLM Workshop on Learning from Situated and Embodied Interaction.
  2. Learning from Synthetic Data Improves Multi-hop Reasoning
    In ICLR (2026). Spotlight talk at NSF-NAIRR annual meeting.
  3. PhantomWiki: On-Demand Datasets for Reasoning and Retrieval Evaluation
    Albert Gong*, Kamilė Stankevičiūtė*, Chao Wan*, Anmol Kabra*, Raphael Thesmar, Johann Lee, Julius Klenke, Carla P. Gomes, and Kilian Q. Weinberger
    In ICML (2025). Oral presentation at ICML Workshop on Long Context Foundation Models.
  4. Exponential Family Model-Based Reinforcement Learning via Score Matching
    In NeurIPS (2022). Oral presentation.
  5. Characterizing the Loss Landscape in Non-Negative Matrix Factorization
    In AAAI (2021).

also known for


(sorted ascending by number of characters per item)
  • biking
  • cooking
  • running
  • juggling
  • being outdoors
  • playing Table Tennis
  • following Formula 1 and motorsports
  • walking fast so that my legs heat up
  • reading books, newspapers, and research papers