--- title: 'Quantum Verifiable Rewards for Post-Training Qiskit Code Assistant' subtitle: 'Training AI models to write better quantum code with quantum hardware verification' summary: Released a new paper on a novel approach to train AI models that can write better quantum code using Qiskit. The approach uses quantum verification at the core, smart training pipeline with DPO and GRPO, and real quantum feedback to ensure generated code works in practice. authors: - juancb tags: - Quantum Computing - AI - Machine Learning - Qiskit - Research - AI for Quantum categories: - Quantum Computing - Artificial Intelligence - Research date: "2025-08-30T00:00:00Z" lastmod: "2025-08-30T00:00:00Z" featured: true draft: false # Featured image # To use, add an image named `featured.jpg/png` to your page's folder. # Placement options: 1 = Full column width, 2 = Out-set, 3 = Screen-width # Focal point options: Smart, Center, TopLeft, Top, TopRight, Left, Right, BottomLeft, Bottom, BottomRight image: placement: 2 caption: 'Quantum Verifiable Rewards for Post-Training Qiskit Code Assistant' focal_point: "Smart" preview_only: false # Projects (optional). # Associate this post with one or more of your projects. # Simply enter your project's folder or file name without extension. # E.g. `projects = ["internal-project"]` references `content/project/deep-learning/index.md`. # Otherwise, set `projects = []`. projects: [] --- 📄 We just released a new paper "Quantum Verifiable Rewards for Post-Training Qiskit Code Assistant" - and it's been getting great attention from folks in the quantum field since day one. 🎉 🔍 What we built: A novel approach to train AI models that can write better quantum code using Qiskit. What makes this interesting: ✅ Quantum verification at the core - Instead of just hoping the AI-generated code works, we actually verify it runs correctly on real quantum systems ✅ Smart training pipeline - We created synthetic quantum problem-test pairs and used both Direct Preference Optimization (DPO) and Group Relative Policy Optimization (GRPO) to align our models ✅ Real quantum feedback - Our models learn directly from quantum hardware and systems rewards, ensuring the code they generate actually works in practice The results: Our best model significantly outperforms existing leading open-source baselines on the challenging Qiskit-HumanEval-hard benchmark. This work presents a solid contribution to making quantum programming more accessible through AI assistance. The approach of using quantum hardware verification to train coding assistants opens up promising directions for future research. Proud of our team's thoughtful work on this project: Nicolas Dupuis, Adarsh Tiwari, Youssef MROUEH, David Kremer, Ismael Faro. The intersection of AI and quantum computing continues to offer compelling research opportunities. Paper: https://arxiv.org/abs/2508.20907 #QuantumComputing #AI #MachineLearning #Qiskit #Research #AIforQuantum --- *Originally shared on [LinkedIn](https://www.linkedin.com/posts/juancb_quantum-verifiable-rewards-for-post-training-activity-7367455154318585857-5ElR) on August 30, 2025 - 76 reactions, 0 comments as of 11/12/2025*