Portrait of Akshay Nambi

Akshay Nambi

Principal Researcher

Connect on LinkedIn

About

I am a Principal Researcher at Microsoft Research India, working at the intersection of AI, machine learning, and systems. My research focuses on Self-Evolving Agentic Systems (SEAs)—AI systems that learn from experience and become more capable, efficient, and reliable over time.

My long-term research goal is trustworthy recursive self-improvement (RSI) for agents. I study controlled learning loops in which models, environments, and agent workflows improve together using task trajectories, feedback, and verified outcomes. Rather than repeatedly solving each task from scratch or relying only on more inference-time computation, these systems turn experience into reusable capability.

The central challenge is deciding what should change and whether the change is genuinely better. A failure may originate in the model, training data, task, environment, verifier, memory, tool, or workflow—or in an interaction among them. My work develops methods to diagnose these failures, create targeted learning experiences, update the appropriate component, and verify that improvements transfer without compromising safety or prior capabilities. I work across both frontier and small language models, developing shared principles and techniques rather than treating them as separate research directions.

Research Areas

1. Self-Evolving Agentic Systems

I build agents that improve through repeated cycles of experience, learning, verification, and redeployment. This includes:

  • co-evolving models, synthetic environments, tasks, data, verifiers, and agent workflows;
  • assigning credit across long-horizon trajectories and interacting system components;
  • learning reusable memories and skills while managing their verification, composition, updating, and retirement; and
  • enabling continual improvement without forgetting, reward hacking, or benchmark overfitting.

Related Works: Echoverse, Echoverse-Web  (opens in new tab)

2. Post-Training for Agentic Models

I develop scalable post-training recipes for models that reason, use tools, operate computers, and recover from failure. My research explores reinforcement learning, dense trajectory-level rewards, rubric-based learning signals, on-policy and iterative distillation, targeted synthetic data, and evolving curricula. The goal is to improve agentic capability efficiently across frontier and small language models

Related work: Fara 1.5 (opens in new tab), ATLAS (opens in new tab), ARTIST (opens in new tab), Agent-Brace (opens in new tab), Self-distillation (opens in new tab), AutoAdapt, (opens in new tab) Think Right (opens in new tab)

3. Trustworthy and Safe Agentic Systems

As agents move from generation to action, ensuring safety, reliability, and alignment becomes critical. My work focuses on building agents that can reason about uncertainty, verify outcomes, and decide when to act or refuse, especially in non-verifiable settings. This includes developing methods for safe tool use, failure detection, and robustness in long-horizon execution, as well as evaluation frameworks that capture real-world risks beyond standard benchmarks.

Related Work: MOSAIC [ICML’26], (opens in new tab) Cascaded SAEs[Neurips’26] (opens in new tab)

4. Population-Scale AI Systems and Copilots

A key focus of my work is translating research into real-world AI systems that operate at scale. I build and deploy agentic copilots for product teams, such as Researcher Agents for deep research and complex workflows, as well as for societal applications in domains like education and agriculture.

Related work: Shiksha Copilot (opens in new tab), MMCT

Please visit my projects and publications page for more details. I have developed and scaled impactful solutions that are actively used by several thousands of users across diverse sectors, including education, agriculture, transportation (opens in new tab), healthcare, and energy.

Internship opportunities (3-6months): I’m always on the lookout for bright students and researchers who have strong hands-on experience in large language models, agentic AI, reinforcement learning, reasoning systems, and scalable ML systems. I particularly value individuals who can move fast and build end-to-end systems. If you are interested in internships or collaborations, please email me your CV and research interests.