Adrián Bazaga

Senior Research Scientist · Tech Lead (SLM) · Post-Training Lead

Microsoft  •  PhD, University of Cambridge

I lead mid-training and post-training for the language models that ship in Windows, spanning synthetic data generation, reinforcement learning (RLHF & RLVR), reward modeling, agentic reasoning, alignment, and evaluation.

Adrián Bazaga, Senior Research Scientist and Post-Training Lead at Microsoft

About me

I’m a Senior Research Scientist, Tech Lead (SLM) and Post-Training Lead at Microsoft, working on large language models (LLMs) and natural language processing (NLP). I teach language models to follow instructions, reason and act more reliably as agents, and stay helpful and safe, and I’m broadly interested in how much capability can be packed into models small enough to run on your own device.

I lead post-training for Aion Instruct, Microsoft’s foundational language model for Windows, unveiled by Satya Nadella at Microsoft Build 2026. I own the end-to-end stack from data preparation → pre-training → mid-training → post-training → evaluation, advancing state-of-the-art foundational language models at global scale. My post-training work spans supervised fine-tuning (SFT), synthetic data generation, direct preference optimization (DPO), reinforcement learning (RLVR and RLHF), reward modeling, and agentic reasoning, driving human alignment across response quality, conciseness, factuality, and safety.

I also led post-training for Mu, Microsoft’s blazing-fast on-device SLM, and built the Windows Settings agent, now live on Copilot+ devices.

I hold a Ph.D. in Machine Learning from the University of Cambridge, advised by Prof. Pietro Liò and Prof. Gos Micklem, with work published at ICLR, ICML, ACL and EMNLP as well as in Nature journals. Previously I interned at Microsoft Research and Amazon AGI, and spent several years at the intersection of artificial intelligence (AI) and biology across startups and research labs. Originally from Spain.

Research

What I work on

All publications →

My research advances the capabilities of foundational language models, spanning mid-training and post-training, synthetic data generation, reinforcement learning (RLHF & RLVR), reward modeling, agentic reasoning, alignment, and efficiency, with broader interests across machine learning such as Geometric & Graph Deep Learning and ML for Science. The works below are a selected snapshot, not a complete picture of everything I’m currently exploring.

1 LLM Training & Data-Efficient Learning

Training more capable language models with less supervision: self-supervised pretraining and distillation, data augmentation that gets more from limited labels, and test-time self-reflection for stronger reasoning.

SFAVEL framework overview

Self-supervised claim verification without labeled data, distilling knowledge from a language model.

TISER timeline self-reflection overview

Timeline self-reflection that improves temporal reasoning with test-time compute; developed during my Amazon Science (AGI) work.

TabMDA method overview

Training-free manifold data augmentation for tabular data via in-context subsetting.

2 Geometric & Graph Deep Learning

Learning on graphs, hypergraphs and richer geometric structure, from sheaf-theoretic message passing to hypergraph-aware language models for text-attributed networks.

HyperBERT model overview

Hypergraph-aware layers fused with language models for node classification.

When and why learnable sheaf Laplacians are necessary in graph neural networks.

3 Machine Learning for Science

Deep learning for scientific discovery: spatiotemporal LLMs for physical simulation, and models for genomics and biomedical imaging that turn raw measurements into insight.

FLUID-LLM overview

Spatiotemporal-aware large language models for computational fluid dynamics.

Genome-wide gene–cancer association discovery overview
Nature Sci. Rep. 2020 Genome-wide Investigation of Gene–Cancer Associations for Target Discovery

Genome-wide discovery of gene–cancer associations to predict therapeutic targets.

Impact

Research shipped to millions

Beyond publications, my work ships in production language models running on Windows devices for millions of users worldwide.

Aion on-device AI on Windows 11

Aion Instruct

Post-training lead for Windows’ foundational language model, unveiled by Satya Nadella at Build 2026 and introduced on the Windows Developer Blog.

Mu encoder–decoder architecture compared to decoder-only

Mu

Post-training lead for Mu, a 0.3B on-device SLM built for blazing-fast, low-memory inference.

News

Latest updates

June 2026
Aion Instruct, Microsoft’s foundational language model for Windows, is here.
  • Unveiled by Satya Nadella at Microsoft Build 2026 and introduced on the Windows Developer Blog.
  • I lead its post-training end to end, spanning SFT, DPO, RLVR and RLHF, reward modeling, and synthetic data generation.
  • Now shipping on Windows devices.
September 2025
Promoted to Senior Research Scientist at Microsoft, co-leading a group developing state-of-the-art Small Language Models, from pre-training to post-training and on-device deployment.
June 2025
January 2025
Joined Microsoft as an AI Research Scientist in London, working on on-device LLM experiences for millions of users worldwide.
September 2024
Our paper HyperBERT was accepted at EMNLP 2024.
August 2024
  • Joined the Amazon Science (AGI) team as a Research Scientist Intern, working on test-time scaling for temporal reasoning with LLMs in Berlin.
  • Joined the reviewer committee for ICLR and ACL.
June 2024

Experience

Where I’ve worked

Senior Research Scientist, Post-Training Lead
Microsoft · Jan 2025 – Present · London, UK

Lead mid-training and post-training for Aion Instruct and Mu: synthetic data generation, SFT, DPO, RLVR, RLHF, reward modeling, and agentic reasoning, shipped to millions of Windows devices.

Research Scientist
Amazon Science (AGI) · Aug – Dec 2024 · Berlin

Led an end-to-end project on test-time scaling for temporal reasoning with LLMs, resulting in a main-track ACL 2025 publication.

Research Scientist
Microsoft Research · May – Aug 2024 · Cambridge, UK

Order-agnostic noise schedules for diffusion models, enabling few-step generation with 2× faster sampling at equal quality.

Education

Ph.D. in Machine Learning
University of Cambridge · 2019 – 2025

Passed without corrections. Thesis on multimodal learning and language models for enhanced knowledge representations. Advised by Pietro Liò & Gos Micklem.

M.Sc. in Artificial Intelligence
Rank 2 / 85 · 2017 – 2019
B.Sc. in Computer Science & Mathematics
Rank 3 / 350 · 2013 – 2017

Honors & Awards

Recognition

  • 2024
    Award for Outstanding PhD Research with Real-World Application
    Cambridge Society for the Application of Research
  • 2024
    Rittner Niederste Hollenberg Award
    University of Cambridge
  • 2020
    Senior Scholarship Award for PhD research
    Fitzwilliam College, University of Cambridge
  • 2019
    Knowledge Transfer Partnership Grant
    UK Research and Innovation
  • 2018
    Selected for the Google Summer of Code ’18 Cohort
    Google

Academic Service & Mentorship

Community

Peer review. Reviewer for ICLR, NeurIPS and ICML, and for the Applied Soft Computing and Expert Systems with Applications journals.

Supervision. Supervisor of PhD and Master’s students at the University of Cambridge, University of Oxford and University of Rome (La Sapienza). Open to collaborations.