Skip to content
#

posttraining

Here are 9 public repositories matching this topic...

Post-training code LLMs: SFT + DPO on preference pairs labelled by a sandbox and static analysis, with controlled reward ablations and reproducible HumanEval evaluation. v2 (in progress): a benchmark of PEFT, alignment and distillation methods for code.

  • Updated Sep 27, 2026
  • Python

Add this topic to your repo

To associate your repository with the posttraining topic, visit your repo's landing page and select "manage topics."

Learn more