← Papers & projects

Undergraduate thesis Thesis

Reinforcement Learning for Legal Reasoning on Multi-Choice QA

Hybrid Zero-RL, distilled-CoT SFT, and GRPO for structured legal reasoning

Reinforcement Learning for Legal Reasoning on Multi-Choice QA
57.6%
Test accuracy
3-stage
Hybrid training pipeline
GRPO
Policy optimization

This project develops a hybrid reinforcement learning framework for structured reasoning on legal multiple-choice question answering.

Training design

  • Designed a Zero-RL → distilled-CoT SFT → GRPO training pipeline.
  • Structured outputs around statute citation and option-by-option analysis.
  • Reached 57.6% accuracy, outperforming SFT-only and RL-only baselines.

Read the thesis for the training design, experimental setup, and ablation results.