Senior Research Scientist · Tech Lead (SLM) · Post-Training Lead
Microsoft • PhD, University of Cambridge
I lead mid-training and post-training for the language models that ship in Windows, spanning synthetic data generation, reinforcement learning (RLHF & RLVR), reward modeling, agentic reasoning, alignment, and evaluation.
I’m a Senior Research Scientist, Tech Lead (SLM) and Post-Training Lead at Microsoft, working on large language models (LLMs) and natural language processing (NLP). I teach language models to follow instructions, reason and act more reliably as agents, and stay helpful and safe, and I’m broadly interested in how much capability can be packed into models small enough to run on your own device.
I lead post-training for Aion Instruct, Microsoft’s foundational language model for Windows, unveiled by Satya Nadella at Microsoft Build 2026. I own the end-to-end stack from data preparation → pre-training → mid-training → post-training → evaluation, advancing state-of-the-art foundational language models at global scale. My post-training work spans supervised fine-tuning (SFT), synthetic data generation, direct preference optimization (DPO), reinforcement learning (RLVR and RLHF), reward modeling, and agentic reasoning, driving human alignment across response quality, conciseness, factuality, and safety.
I also led post-training for Mu, Microsoft’s blazing-fast on-device SLM, and built the Windows Settings agent, now live on Copilot+ devices.
I hold a Ph.D. in Machine Learning from the University of Cambridge, advised by Prof. Pietro Liò and Prof. Gos Micklem, with work published at ICLR, ICML, ACL and EMNLP as well as in Nature journals. Previously I interned at Microsoft Research and Amazon AGI, and spent several years at the intersection of artificial intelligence (AI) and biology across startups and research labs. Originally from Spain.
Research
My research advances the capabilities of foundational language models, spanning mid-training and post-training, synthetic data generation, reinforcement learning (RLHF & RLVR), reward modeling, agentic reasoning, alignment, and efficiency, with broader interests across machine learning such as Geometric & Graph Deep Learning and ML for Science. The works below are a selected snapshot, not a complete picture of everything I’m currently exploring.
Training more capable language models with less supervision: self-supervised pretraining and distillation, data augmentation that gets more from limited labels, and test-time self-reflection for stronger reasoning.
Self-supervised claim verification without labeled data, distilling knowledge from a language model.
Timeline self-reflection that improves temporal reasoning with test-time compute; developed during my Amazon Science (AGI) work.
Learning on graphs, hypergraphs and richer geometric structure, from sheaf-theoretic message passing to hypergraph-aware language models for text-attributed networks.
Impact
Beyond publications, my work ships in production language models running on Windows devices for millions of users worldwide.

Post-training lead for Windows’ foundational language model, unveiled by Satya Nadella at Build 2026 and introduced on the Windows Developer Blog.
News
Experience
Lead mid-training and post-training for Aion Instruct and Mu: synthetic data generation, SFT, DPO, RLVR, RLHF, reward modeling, and agentic reasoning, shipped to millions of Windows devices.
Led an end-to-end project on test-time scaling for temporal reasoning with LLMs, resulting in a main-track ACL 2025 publication.
Order-agnostic noise schedules for diffusion models, enabling few-step generation with 2× faster sampling at equal quality.
Passed without corrections. Thesis on multimodal learning and language models for enhanced knowledge representations. Advised by Pietro Liò & Gos Micklem.
Honors & Awards
Academic Service & Mentorship
Peer review. Reviewer for ICLR, NeurIPS and ICML, and for the Applied Soft Computing and Expert Systems with Applications journals.
Supervision. Supervisor of PhD and Master’s students at the University of Cambridge, University of Oxford and University of Rome (La Sapienza). Open to collaborations.