Summary

AI Researcher/Engineer focused on LLM post-training, alignment, logical reasoning, and agentic RL. Ph.D. graduate from the University of Auckland with hands-on experience in SFT/DPO/RLHF-style alignment, reward-model design, NLI-as-reward, multimodal/document AI, and production-oriented evaluation loops. Recent work includes Hybrid-DPO for balancing logical grounding and fluency, Conflict-Aware Fusion for mitigating Logic Inertia, and applied Agent/FinTech/remote-sensing systems.

Core Skills

  • Post-Training / Alignment: SFT, DPO, Hybrid-DPO, RLHF/PPO, GRPO, Reward Model design, NLI-as-reward, preference data synthesis.
  • Reasoning / Data: logical reasoning, AMR, data augmentation, OOD generalisation, Chain-of-Thought, multi-turn decision trajectory synthesis.
  • Training Engineering: PyTorch, PEFT/LoRA, Flash-Attention 2, vLLM, Megatron, int4 quantisation, single-GPU 7B-class training workflows.
  • Agentic / Simulation: LLM Agents, Tool Use, RL simulation environments, PPO reward alignment, financial AI Agents, multi-step execution traces.
  • Multimodal / Document AI: Qwen3.5-9B, Gemma 4, InternVL2, LayoutLMv3, ERNIE-LayoutX, OCR, long-sequence extension to 4096.

Work & Project Experience

Large Language Model Post-Training, Alignment and Logical Reasoning (Ph.D.) Strong AI Lab / NAO Institute, UoA, Auckland
Research Project Leader / Developer 02/20 – 09/25
  • RLearner-LLM / Hybrid-DPO: designed an automated preference pipeline combining DeBERTa-v3 NLI signals and Verifier LLM rewards to reduce verbosity bias in standard preference signals; validated across LLaMA-2-13B, Qwen3-8B and Gemma 4 families, with up to 6× NLI-alignment improvement and 95% pairwise win rate for the Qwen3-8B version.
  • Conflict-Aware Fusion / Logic Inertia: defined and quantified Logic Inertia, a failure mode where LLM reasoning collapses under conflicting rules; proposed a dual-process cognitive architecture separating premise verification from logical derivation, reaching 100% accuracy on contradiction-injection and base-task stress tests.
  • AMR-LDA: proposed AMR-based logic-equivalent data augmentation with GPT-4 prompt augmentation; achieved #1 on the ReClor leaderboard and became the first team above 90% on the hidden test set. paper, source code, model weights.
  • OOD logical reasoning evaluation: designed four structural perturbation stress tests including rule deletion, contradiction injection, logic-preserving rewriting and multi-rule equivalence; revealed significant robustness gaps in generative and discriminative LLMs, with related IJCAI 2024 work cited 100+ times.
  • Iterative Enhancement / Educational LLM: developed a closed-loop explanation generation and evaluation framework for learnersourced multiple-choice questions, accepted by AAAI 2025 and AGI@ICLR 2024. paper, source code.
  • OpenAI Evals contributions: openai/evals#648, openai/evals#651.
Xtracta — Multimodal LLM/VLM Training, Evaluation and Document AI Xtracta, Auckland, New Zealand
Artificial Intelligence Researcher / Engineer 07/22 – Now
  • Multimodal model training and evaluation: continually trained and evaluated Qwen3.5-9B, Gemma 4 E4B/31B and InternVL2 with PEFT/LoRA for document and privacy-related benchmarks including CustodianAI medical PII redaction.
  • Integrated Flash-Attention 2 to reduce FP16 peak memory by up to 50%; used int4 quantisation to enable complete training workflows on a single A4090 24GB setup for targeted models.
  • Long-sequence extension: extended LayoutLMv3 / ERNIE-LayoutX maximum sequence length from 512 to 4096 via sliding-window attention and Longformer-style global attention masks, improving F1 on XFUND, FUNSD and internal token-classification / relation-extraction datasets without obvious GPU-memory growth.
  • LayoutLMv3 pretraining reproduction: reproduced the non-open-sourced MLM + MIM + WPA multimodal pretraining process and integrated global-attention masks for stronger text–vision alignment.
  • Applied affine-transform augmentation to improve line-alignment robustness for document extraction; successfully supported Callaghan Innovation RDTI applications for 2022/2023.
NLP R&D Internships & Medical NLP AIIT, Peking University / Precision Driven Health
R&D Engineer / Research Project Leader 2018 – 2020
  • Developed a meeting-robot NLP pipeline for abstract extraction, text segmentation, topic prediction and multi-turn QA, with reusable APIs for meeting-record processing.
  • Proposed HBAM for medical text similarity, published at ACSW 2020 with 90+ citations; built a medical PII detection and redaction pipeline and privacy-preserving LLM toolkit that supported the establishment of CustodianAI and AUT Venture investment of NZD 60,000. Demo.

Education

University of Auckland (UoA)Auckland, New Zealand
Ph.D. in Computer Science, fully funded; supervised by Prof. Michael Witbrock and Assoc Prof. Jiamou Liu 02/20 – 09/25
University of Auckland (UoA)Auckland, New Zealand
Bachelor of Science (Honours) in Computer Science (First Class), GPA: 7/9 07/18 – 09/19
  • Precision Driven Health & Orion Health Summer Research Scholarship

Paper List

  • Qiming Bao, Juho Leinonen, Alex Yuxuan Peng, Wanjun Zhong, Tim Pistotti, Alice Huang, Paul Denny, Michael Witbrock, Jiamou Liu. Exploring Iterative Enhancement for Improving Learnersourced Multiple-Choice Question Explanations with Large Language Models, Proceedings of the AAAI Conference on Artificial Intelligence (2025). https://ojs.aaai.org/index.php/AAAI/article/view/35164
  • Qiming Bao, Alex Peng, Zhenyun Deng, Wanjun Zhong, Gaël Gendron, Neşet Tan, Nathan Young, Yang Chen, Yonghua Zhu, Michael Witbrock, Jiamou Liu. Abstract Meaning Representation-Based Logic-Driven Data Augmentation for Logical Reasoning., The Findings of ACL (2024). https://doi.org/10.18653/v1/2024.findings-acl.353
  • Qiming Bao, Gaël Gendron, Alex Peng, Neset Tan, Michael Witbrock, Jiamou Liu. Assessing and Enhancing the Robustness of Large Language Models with Task Structure Variations for Logical Reasoning., ICONIP (2024). https://doi.org/10.48550/arXiv.2310.09430
  • Qiming Bao, Alex Peng, Tim Hartill, Neset Tan, Zhenyun Deng, Michael Witbrock, Jiamou Liu. Multi-Step Deductive Reasoning Over Natural Language: An Empirical Study on Out-of-Distribution Generalisation, IJCLR-NeSy (2022). https://ceur-ws.org/Vol-3212/paper15.pdf
  • Nathan Young, Qiming Bao, Joshua Ljudo Bensemann, Michael J. Witbrock. AbductionRules: Training Transformers to Explain Unexpected Inputs, The Findings of ACL (2022). https://doi.org/10.18653/v1/2022.findings-acl.19
  • Gaël Gendron, Qiming Bao, Michael Witbrock, Gillian Dobbie. Large Language Models Are Not Strong Abstract Reasoners, IJCAI (2024). https://www.ijcai.org/proceedings/2024/693
  • Lin Ni, Qiming Bao, Xiaoxuan Li, Qianqian Qi, Paul Denny, Jim Warren, Michael Witbrock, Jiamou Liu. DeepQR: Neural-based Quality Ratings for Learnersourced Multiple-Choice Questions, Proceedings of the AAAI Conference on Artificial Intelligence (2022). https://doi.org/10.1609/aaai.v36i11.21562
  • Qianqian Qi, Qiming Bao*, Alex Yuxuan Peng, Jiamou Liu, Michael Witbrock. A Dynamic Prompt-tuning Method for Data Augmentation with Associated Knowledge, ICLR TinyPapers (2023). https://openreview.net/pdf?id=hli7A0ioiS_
  • Qiming Bao, Lin Ni, Jiamou Liu. HHH: An Online Medical Chatbot System based on Knowledge Graph and Hierarchical Bi-Directional Attention, ACSW (2020). https://doi.org/10.1145/3373017.3373049
  • Zhongsheng Wang, Jiamou Liu, Qiming Bao, Hongfei Rong, Jingfeng Zhang. ChatLogic: Integrating Logic Programming with Large Language Models for Multi-step Reasoning, NucLeaR@AAAI (2024). https://doi.org/10.48550/arXiv.2407.10162
  • Neset TAN, Trung Nguyen, Josh Bensemann, Alex Peng, Qiming Bao, Yang Chen, Mark Gahegan, Michael Witbrock. Multi2Claim: Generating Scientific Claims from Multi-Choice Questions for Scientific Fact-Checking, EACL (2023). https://doi.org/10.18653/v1/2023.eacl-main.194
  • Neset TAN, Alex Peng, Joshua Bensemann, Qiming Bao, Tim Hartill, Mark Gahegan, Michael Witbrock. Input-length-shortening and text generation via attention values, AAAI-EMC^2 (2023). https://doi.org/10.48550/arXiv.2303.07585
  • Invited Speaker/Visiting Scholar

  • Microsoft Research Asia Invited Talk 2022 invited by Dr. Bei Chen (Invitation Letter) (Presentation Slide) (Recording)
  • Samsung AI Center Cambridge UK Invited Talk 2022 invited by Dr. Cristina Cornelio(Invitation Letter) (Presentation Slide) (Recording)
  • IEEE Vehicular Technology Society (VTS) New Zealand North Chapter and IEEE New Zealand North Section SIGHT Group 2022 invited by Dr. William Liu (Invitation Letter) (Presentation Slide) (Recording)
  • ZJU-NLP Group, Zhejiang University 2023 invited by Prof. Huajun Chen and Prof. Ningyu Zhang
  • NLP Group, The University of Melbourne Invited Talk 2023 invited by Mr. Ming-Bin Chen and Dr. Jey Han Lau (Invitation Letter) (Presentation Slide)
  • Institute of Automation, Chinese Academy of Sciences Invited Talk 2023 (Invitation Letter) (Presentation Slide)
  • University of Massachusetts - Amherst Invited Talk 2024 invited by Prof. Andrew Lan (Invitation Letter) (Presentation Slide)
  • Penn State University & University of Auckland Online Workshop 2024 Day 1 Session 2 Children's Future, Intercultural Learning (Invitation Letter) (Presentation Slide) (Recording)
  • Logic and AI Seminar 2025 (Tsinghua University & Peking University) invited by Prof. Fenrong Liu and A/Prof. Haoxuan Li (Invitation Letter)
  • Max Planck Institute For Software Systems invited by Prof. Adish Singla (Invitation Letter) (Presentation Slide)
  • Technical University of Munich invited by Dr. Stefan Fuchs and Prof. Dr.-Ing. André Borrmann (Invitation Letter)