Summary
AI Researcher/Engineer focused on LLM post-training, alignment, logical reasoning, and agentic RL. Ph.D. graduate from the University of Auckland with hands-on experience in SFT/DPO/RLHF-style alignment, reward-model design, NLI-as-reward, multimodal/document AI, and production-oriented evaluation loops. Recent work includes Hybrid-DPO for balancing logical grounding and fluency, Conflict-Aware Fusion for mitigating Logic Inertia, and applied Agent/FinTech/remote-sensing systems.
Core Skills
- Post-Training / Alignment: SFT, DPO, Hybrid-DPO, RLHF/PPO, GRPO, Reward Model design, NLI-as-reward, preference data synthesis.
- Reasoning / Data: logical reasoning, AMR, data augmentation, OOD generalisation, Chain-of-Thought, multi-turn decision trajectory synthesis.
- Training Engineering: PyTorch, PEFT/LoRA, Flash-Attention 2, vLLM, Megatron, int4 quantisation, single-GPU 7B-class training workflows.
- Agentic / Simulation: LLM Agents, Tool Use, RL simulation environments, PPO reward alignment, financial AI Agents, multi-step execution traces.
- Multimodal / Document AI: Qwen3.5-9B, Gemma 4, InternVL2, LayoutLMv3, ERNIE-LayoutX, OCR, long-sequence extension to 4096.
Work & Project Experience
Large Language Model Post-Training, Alignment and Logical Reasoning (Ph.D.)
Strong AI Lab / NAO Institute, UoA, Auckland
Research Project Leader / Developer 02/20 – 09/25
Research Project Leader / Developer 02/20 – 09/25
- RLearner-LLM / Hybrid-DPO: designed an automated preference pipeline combining DeBERTa-v3 NLI signals and Verifier LLM rewards to reduce verbosity bias in standard preference signals; validated across LLaMA-2-13B, Qwen3-8B and Gemma 4 families, with up to 6× NLI-alignment improvement and 95% pairwise win rate for the Qwen3-8B version.
- Conflict-Aware Fusion / Logic Inertia: defined and quantified Logic Inertia, a failure mode where LLM reasoning collapses under conflicting rules; proposed a dual-process cognitive architecture separating premise verification from logical derivation, reaching 100% accuracy on contradiction-injection and base-task stress tests.
- AMR-LDA: proposed AMR-based logic-equivalent data augmentation with GPT-4 prompt augmentation; achieved #1 on the ReClor leaderboard and became the first team above 90% on the hidden test set. paper, source code, model weights.
- OOD logical reasoning evaluation: designed four structural perturbation stress tests including rule deletion, contradiction injection, logic-preserving rewriting and multi-rule equivalence; revealed significant robustness gaps in generative and discriminative LLMs, with related IJCAI 2024 work cited 100+ times.
- Iterative Enhancement / Educational LLM: developed a closed-loop explanation generation and evaluation framework for learnersourced multiple-choice questions, accepted by AAAI 2025 and AGI@ICLR 2024. paper, source code.
- OpenAI Evals contributions: openai/evals#648, openai/evals#651.
Xtracta — Multimodal LLM/VLM Training, Evaluation and Document AI
Xtracta, Auckland, New Zealand
Artificial Intelligence Researcher / Engineer 07/22 – Now
Artificial Intelligence Researcher / Engineer 07/22 – Now
- Multimodal model training and evaluation: continually trained and evaluated Qwen3.5-9B, Gemma 4 E4B/31B and InternVL2 with PEFT/LoRA for document and privacy-related benchmarks including CustodianAI medical PII redaction.
- Integrated Flash-Attention 2 to reduce FP16 peak memory by up to 50%; used int4 quantisation to enable complete training workflows on a single A4090 24GB setup for targeted models.
- Long-sequence extension: extended LayoutLMv3 / ERNIE-LayoutX maximum sequence length from 512 to 4096 via sliding-window attention and Longformer-style global attention masks, improving F1 on XFUND, FUNSD and internal token-classification / relation-extraction datasets without obvious GPU-memory growth.
- LayoutLMv3 pretraining reproduction: reproduced the non-open-sourced MLM + MIM + WPA multimodal pretraining process and integrated global-attention masks for stronger text–vision alignment.
- Applied affine-transform augmentation to improve line-alignment robustness for document extraction; successfully supported Callaghan Innovation RDTI applications for 2022/2023.
NLP R&D Internships & Medical NLP
AIIT, Peking University / Precision Driven Health
R&D Engineer / Research Project Leader 2018 – 2020
R&D Engineer / Research Project Leader 2018 – 2020
- Developed a meeting-robot NLP pipeline for abstract extraction, text segmentation, topic prediction and multi-turn QA, with reusable APIs for meeting-record processing.
- Proposed HBAM for medical text similarity, published at ACSW 2020 with 90+ citations; built a medical PII detection and redaction pipeline and privacy-preserving LLM toolkit that supported the establishment of CustodianAI and AUT Venture investment of NZD 60,000. Demo.
Education
University of Auckland (UoA)Auckland, New Zealand
Ph.D. in Computer Science, fully funded; supervised by Prof. Michael Witbrock and Assoc Prof. Jiamou Liu 02/20 – 09/25
Ph.D. in Computer Science, fully funded; supervised by Prof. Michael Witbrock and Assoc Prof. Jiamou Liu 02/20 – 09/25
- DAAD AINeT Fellow 2025 on Natural Language Processing
- Graduate Teaching Assistant / Research Assistant
University of Auckland (UoA)Auckland, New Zealand
Bachelor of Science (Honours) in Computer Science (First Class), GPA: 7/9 07/18 – 09/19
Bachelor of Science (Honours) in Computer Science (First Class), GPA: 7/9 07/18 – 09/19
- Precision Driven Health & Orion Health Summer Research Scholarship