Adrián Bazaga

Senior Researcher · Tech Lead (SLM) · Post-Training Lead

Microsoft  •  PhD, University of Cambridge

I lead post-training for the language models that ship in Windows, spanning data, reinforcement learning, alignment, and evaluation.

Adrián Bazaga

About me

I’m a Senior Researcher, Tech Lead (SLM) and Post-Training Lead at Microsoft. I teach language models to follow instructions, reason more reliably, and stay helpful and safe, and I’m broadly interested in how much capability can be packed into models small enough to run on your own device.

I lead post-training for Aion Instruct, Microsoft’s foundational language model for Windows, unveiled by Satya Nadella at Microsoft Build 2026. I own the end-to-end stack from data preparation → pre-training → mid-training → post-training → evaluation, advancing state-of-the-art foundational language models at global scale. My post-training work spans supervised fine-tuning (SFT), synthetic data generation, direct preference optimization (DPO), reinforcement learning (RLVR and RLHF), and reward modeling, driving human alignment across response quality, conciseness, factuality, and safety.

I also led post-training for Mu, Microsoft’s blazing-fast on-device SLM, and built the Windows Settings agent, now live on Copilot+ devices.

I hold a Ph.D. in Machine Learning from the University of Cambridge, advised by Prof. Pietro Liò and Prof. Gos Micklem, with work published at ICLR, ICML, ACL and EMNLP as well as in Nature journals. Previously I interned at Microsoft Research and Amazon AGI, and spent several years at the intersection of AI and biology across startups and research labs. Originally from Spain.

Research

What I work on

All publications →

My research advances the capabilities of foundational language models, spanning post-training, alignment, reasoning, and efficiency, with broader interests across machine learning such as Geometric & Graph Deep Learning and ML for Science. The works below are a selected snapshot, not a complete picture of everything I’m currently exploring.

1 LLM Training & Data-Efficient Learning

Training more capable language models with less supervision: self-supervised pretraining and distillation, data augmentation that gets more from limited labels, and test-time self-reflection for stronger reasoning.

SFAVEL framework overview

Self-supervised claim verification without labeled data, distilling knowledge from a language model.

TISER timeline self-reflection overview

Timeline self-reflection that improves temporal reasoning with test-time compute; developed during my Amazon Science (AGI) work.

TabMDA method overview

Training-free manifold data augmentation for tabular data via in-context subsetting.

2 Geometric & Graph Deep Learning

Learning on graphs, hypergraphs and richer geometric structure, from sheaf-theoretic message passing to hypergraph-aware language models for text-attributed networks.

HyperBERT model overview

Hypergraph-aware layers fused with language models for node classification.

When and why learnable sheaf Laplacians are necessary in graph neural networks.

3 Machine Learning for Science

Deep learning for scientific discovery: spatiotemporal LLMs for physical simulation, and models for genomics and biomedical imaging that turn raw measurements into insight.

FLUID-LLM overview

Spatiotemporal-aware large language models for computational fluid dynamics.

Genome-wide gene–cancer association discovery overview
Nature Sci. Rep. 2020 Genome-wide Investigation of Gene–Cancer Associations for Target Discovery

Genome-wide discovery of gene–cancer associations to predict therapeutic targets.

Impact

Research shipped to millions

Beyond publications, my work ships in production language models running on Windows devices for millions of users worldwide.

Aion on-device AI on Windows 11

Aion Instruct

Post-training lead for Windows’ foundational language model, unveiled by Satya Nadella at Build 2026 and introduced on the Windows Developer Blog.

Mu encoder–decoder architecture compared to decoder-only

Mu

Post-training lead for Mu, a 0.3B on-device SLM built for blazing-fast, low-memory inference.

News

Latest updates

Experience

Where I’ve worked

Senior Researcher, Post-Training Lead
Microsoft · Jan 2025 – Present · London, UK

Lead post-training for Aion Instruct and Mu: SFT, DPO, RLVR, RLHF, reward modeling, and synthetic data generation, shipped to millions of Windows devices.

Research Scientist
Amazon Science (AGI) · Aug – Dec 2024 · Berlin

Led an end-to-end project on test-time scaling for temporal reasoning with LLMs, resulting in a main-track ACL 2025 publication.

Research Scientist
Microsoft Research · May – Aug 2024 · Cambridge, UK

Order-agnostic noise schedules for diffusion models, enabling few-step generation with 2× faster sampling at equal quality.

Education

Ph.D. in Machine Learning
University of Cambridge · 2019 – 2025

Passed without corrections. Thesis on multimodal learning and language models for enhanced knowledge representations. Advised by Pietro Liò & Gos Micklem.

M.Sc. in Artificial Intelligence
Rank 2 / 85 · 2017 – 2019
B.Sc. in Computer Science & Mathematics
Rank 3 / 350 · 2013 – 2017

Honors & Awards

Recognition

  • 2024
    Award for Outstanding PhD Research with Real-World Application
    Cambridge Society for the Application of Research
  • 2024
    Rittner Niederste Hollenberg Award
    University of Cambridge
  • 2020
    Senior Scholarship Award for PhD research
    Fitzwilliam College, University of Cambridge
  • 2019
    Knowledge Transfer Partnership Grant
    UK Research and Innovation
  • 2018
    Selected for the Google Summer of Code ’18 Cohort
    Google

Academic Service & Mentorship

Community

Peer review. Reviewer for ICLR, NeurIPS and ICML, and for the Applied Soft Computing and Expert Systems with Applications journals.

Supervision. Supervisor of PhD and Master’s students at the University of Cambridge, University of Oxford and University of Rome (La Sapienza). Open to collaborations.