Publication
Diffeomorphic Optimization
Video
Reinforce Adjoint Matching: Scaling Diffusion RL
Diffusion and flow-matching models scale because pretraining is supervised regression: a clean sample is noised analytically, and a model regresses against a closed-form target. RL post-training aligns the model with a reward. In image generation,…
Microsoft Research Blog
SkillOpt: Agent skills as trainable parameters
AI agents often fail because their instructions, or skills, are manually modified with no guarantee of improvement. Learn how SkillOpt turns skill editing into a training process, making agent behavior more reliable without changing model…