Senior Researcher · Tech Lead (SLM) · Post-Training Lead
Microsoft • PhD, University of Cambridge
I lead post-training for the language models that ship in Windows, spanning data, reinforcement learning, alignment, and evaluation.
I’m a Senior Researcher, Tech Lead (SLM) and Post-Training Lead at Microsoft. I teach language models to follow instructions, reason more reliably, and stay helpful and safe, and I’m broadly interested in how much capability can be packed into models small enough to run on your own device.
I lead post-training for Aion Instruct, Microsoft’s foundational language model for Windows, unveiled by Satya Nadella at Microsoft Build 2026. I own the end-to-end stack from data preparation → pre-training → mid-training → post-training → evaluation, advancing state-of-the-art foundational language models at global scale. My post-training work spans supervised fine-tuning (SFT), synthetic data generation, direct preference optimization (DPO), reinforcement learning (RLVR and RLHF), and reward modeling, driving human alignment across response quality, conciseness, factuality, and safety.
I also led post-training for Mu, Microsoft’s blazing-fast on-device SLM, and built the Windows Settings agent, now live on Copilot+ devices.
I hold a Ph.D. in Machine Learning from the University of Cambridge, advised by Prof. Pietro Liò and Prof. Gos Micklem, with work published at ICLR, ICML, ACL and EMNLP as well as in Nature journals. Previously I interned at Microsoft Research and Amazon AGI, and spent several years at the intersection of AI and biology across startups and research labs. Originally from Spain.
Research
My research advances the capabilities of foundational language models, spanning post-training, alignment, reasoning, and efficiency, with broader interests across machine learning such as Geometric & Graph Deep Learning and ML for Science. The works below are a selected snapshot, not a complete picture of everything I’m currently exploring.
Training more capable language models with less supervision: self-supervised pretraining and distillation, data augmentation that gets more from limited labels, and test-time self-reflection for stronger reasoning.
Self-supervised claim verification without labeled data, distilling knowledge from a language model.
Timeline self-reflection that improves temporal reasoning with test-time compute; developed during my Amazon Science (AGI) work.
Learning on graphs, hypergraphs and richer geometric structure, from sheaf-theoretic message passing to hypergraph-aware language models for text-attributed networks.
Impact
Beyond publications, my work ships in production language models running on Windows devices for millions of users worldwide.

Post-training lead for Windows’ foundational language model, unveiled by Satya Nadella at Build 2026 and introduced on the Windows Developer Blog.
News
Experience
Lead post-training for Aion Instruct and Mu: SFT, DPO, RLVR, RLHF, reward modeling, and synthetic data generation, shipped to millions of Windows devices.
Led an end-to-end project on test-time scaling for temporal reasoning with LLMs, resulting in a main-track ACL 2025 publication.
Order-agnostic noise schedules for diffusion models, enabling few-step generation with 2× faster sampling at equal quality.
Education
Passed without corrections. Thesis on multimodal learning and language models for enhanced knowledge representations. Advised by Pietro Liò & Gos Micklem.
Honors & Awards
Academic Service & Mentorship
Peer review. Reviewer for ICLR, NeurIPS and ICML, and for the Applied Soft Computing and Expert Systems with Applications journals.
Supervision. Supervisor of PhD and Master’s students at the University of Cambridge, University of Oxford and University of Rome (La Sapienza). Open to collaborations.