Adrián Bazaga's Homepage
Senior Researcher, Tech Lead (SLM) & Post-Training Lead @ Microsoft
Welcome to my website!
Background:
Hi, I’m Adrián, a Senior Researcher, Tech Lead (SLM) and Post-Training Lead at Microsoft. I’m an established researcher with significant experience in model pre-training, post-training, reasoning, test-time scaling, tool-use, alignment, agents, LLM research, as well as reinforcement learning. I lead post-training for Aion Instruct, Microsoft’s foundational language model for Windows, unveiled by Satya Nadella at Microsoft Build 2026 and shipping on Windows devices, owning the end-to-end stack from data preparation → pretraining → mid-training → post-training → evaluation, and shipping agentic experiences directly on user devices for millions of users at global scale. My post-training work spans supervised fine-tuning (SFT), direct preference optimization (DPO), reinforcement learning (RL), verifiable rewards, reward modeling, and agentic tool-use, driving human alignment across response quality, conciseness, factuality, and safety.
I led post-training for Mu, Microsoft’s blazing-fast on-device SLM, driving the design and implementation of its training pipelines. I also led the development of the Windows Settings AI agent, already live on Windows Copilot+ devices. These efforts are part of my broader vision to enable seamless, deeply integrated AI experiences for everyone.
I hold a Ph.D. in Machine Learning from the University of Cambridge, where I conducted research under the supervision of Prof. Pietro Liò and Prof. Gos Micklem. My work has been published in leading Machine Learning conferences such as ICLR, ICML, ACL and EMNLP, as well as in Nature journals. Previously, I gained research experience through internships at Microsoft Research and Amazon AGI, where I explored novel training schemes to enhance few-step generation in diffusion models, and test-time scaling for temporal reasoning with Large Language Models (LLMs). Prior to that, I spent ~5 years in various startups, working at the intersection of AI and biology.
Research Interests:
My research focuses on advancing the capabilities of AI in the areas of foundational LLMs, reasoning, tool usage, and multimodality. Currently, I’m broadly interested in devising efficient SLM architectures, inventing data-efficient optimization techniques, advancing post-training with reinforcement learning (RL), verifiable rewards, and reward modeling, expanding agentic tool-use, and ensuring human alignment across response quality, conciseness, factuality, and safety.
Beyond Research:
In addition to my core research, I’d like to explore how generative models can improve education and governance. If you’re working on high-impact, real-world deployments in these areas, I’m always open to collaborate 👐.
News
| Jun 2, 2026 | Aion Instruct, Microsoft’s foundational language model for Windows, was unveiled by Satya Nadella at Microsoft Build 2026. I lead its post-training, now shipping on Windows devices. 🚀 |
|---|---|
| Sep 1, 2025 | I have been promoted to Senior Research Scientist at Microsoft, now co-leading a group developing state-of-the-art Small Language Models (SLMs), from pre-training to post-training and all the way to on-device deployment. ⭐ |
| Jun 23, 2025 | We have launched Mu, our 0.3B ‘micro-size’ language model, built for blazing-fast on-device inference and already powering native agentic experiences on Windows devices. 🚀 |
| Jun 1, 2025 | [Paper] Our paper “Learning to Reason Over Time: Timeline Self-Reflection for Improved Temporal Reasoning in Language Models” has been accepted at ACL 2025 (Main) 🎉 |
| Jan 3, 2025 | I joined Microsoft as an AI Research Scientist in London (UK). Excited to work on delivering on-device LLM-based AI experiences for millions of users worldwide. ⭐ |
| Sep 20, 2024 | [Paper] Our paper “HyperBERT: Mixing Hypergraph-Aware Layers with Language Models for Node Classification on Text-Attributed Hypergraphs” has been accepted at EMNLP 2024 🎉 |
| Aug 15, 2024 | I joined Amazon Science AGI team as a Research Scientist Intern to work on test-time scaling for temporal reasoning with LLMs alongside Bill Byrne, Rexhina Blloshmi and Adrià de Gispert, in Berlin (Germany). ⭐ |
| Aug 13, 2024 | I’m now part of the Reviewer Committee for the International Conference on Learning Representations (ICLR) and ACL conferences. 👍 |
| Jun 16, 2024 | Our paper “FLUID-LLM: Learning Computational Fluid Dynamics with Spatiotemporal-aware Large Language Models” is now on arXiv. 📋 |
| Jun 5, 2024 | [Paper] Our paper “TabMDA: Tabular Manifold Data Augmentation for Any Classifier using Transformers with In-context Subsetting” has been accepted at ICML 2024 🎉 |
Selected Publications
- ICLROn the Necessity of Learnable Sheaf LaplaciansIn ICLR 2026 (International Conference on Learning Representations) 2026