Principal Research Scientist · Tech Lead (SLM) · Post-Training Lead
Microsoft • PhD, University of Cambridge
I lead end-to-end post-training for the Aion Instruct model family, Microsoft’s foundational language models for Windows, spanning SFT, DPO, RLVR and RLHF, reward modeling, agentic tool use, safety, alignment, evaluation, and release readiness.
I’m a Principal Research Scientist, Tech Lead (SLM) and Post-Training Lead at Microsoft, working on large language models (LLMs) and natural language processing (NLP). I teach language models to follow instructions, reason and act more reliably as agents, and stay helpful and safe, and I’m broadly interested in how much capability can be packed into models small enough to run on your own device.
I lead end-to-end post-training for the Aion Instruct model family, Microsoft’s foundational language models for Windows, presented by Satya Nadella at Microsoft Build 2026 and shipped on Windows devices. My work spans supervised fine-tuning (SFT), direct preference optimization (DPO), reinforcement learning with verifiable and human-feedback rewards (RLVR and RLHF), reward modeling, agentic tool use, safety and alignment, data, evaluation, and release readiness.
I also led post-training for Mu, Microsoft’s blazing-fast on-device SLM, and built the Windows Settings agent, now live on Copilot+ devices.
I hold a Ph.D. in Machine Learning from the University of Cambridge, advised by Prof. Pietro Liò and Prof. Gos Micklem, with work published at ICLR, ICML, ACL and EMNLP as well as in Nature journals. Previously I interned at Microsoft Research and Amazon AGI, and spent several years at the intersection of artificial intelligence (AI) and biology across startups and research labs. Originally from Spain.
Research
My research advances the capabilities of foundational language models, spanning mid-training and post-training, synthetic data generation, reinforcement learning (RLHF & RLVR), reward modeling, agentic reasoning, alignment, and efficiency, with broader interests across machine learning such as Geometric & Graph Deep Learning and ML for Science. The works below are a selected snapshot, not a complete picture of everything I’m currently exploring.
Training more capable language models with less supervision: self-supervised pretraining and distillation, data augmentation that gets more from limited labels, and test-time self-reflection for stronger reasoning.
Self-supervised claim verification without labeled data, distilling knowledge from a language model.
Timeline self-reflection that improves temporal reasoning with test-time compute; developed during my Amazon Science (AGI) work.
Learning on graphs, hypergraphs and richer geometric structure, from sheaf-theoretic message passing to hypergraph-aware language models for text-attributed networks.
Impact
Beyond publications, my work ships in production language models running on Windows devices for millions of users worldwide.

Post-training lead for Windows’ foundational language model family, unveiled by Satya Nadella at Build 2026 and introduced on the Windows Developer Blog.
News
Experience
Lead end-to-end post-training and the model roadmap for the Aion Instruct model family across data, SFT, DPO, RLVR and RLHF, reward modeling, agentic tool use, safety, and evaluation; also led post-training for Mu, shipping models and agentic experiences to millions of Windows devices.
Led an end-to-end project on test-time scaling for temporal reasoning with LLMs, resulting in a main-track ACL 2025 publication.
Order-agnostic noise schedules for diffusion models, enabling few-step generation with 2× faster sampling at equal quality.
Passed without corrections. Thesis on multimodal learning and language models for enhanced knowledge representations. Advised by Pietro Liò & Gos Micklem.
Honors & Awards
Academic Service & Mentorship
Peer review. Reviewer for ICLR, NeurIPS and ICML, and for the Applied Soft Computing and Expert Systems with Applications journals.
Supervision. Supervisor of PhD and Master’s students at the University of Cambridge, University of Oxford and University of Rome (La Sapienza). Open to collaborations.