dpo
Here are 479 public repositories matching this topic...
A modular, explainable Retrieval-Augmented Generation (RAG) pipeline for automated fact-checking on the FEVER dataset. Includes multi-model evaluation (Mistral, Qwen, GPT), PEFT & DPO fine-tuning, and token-level explainability with Captum, SHAP, and LIME
-
Updated
Oct 13, 2025 - Jupyter Notebook
Experiments with SFT, LoRA, QLoRA, RL post-training
-
Updated
Jun 22, 2026 - Jupyter Notebook
From-scratch RLHF pipeline: reward modeling + PPO (plus DPO & GRPO) in PyTorch — with GAE, KL control, LoRA, multi-GPU, and LLM-as-judge eval.
-
Updated
Jul 6, 2026 - Python
SFT + DPO fine-tuning of Qwen3-1.7B for math reasoning (GSM8K/MATH), QLoRA on a single 6GB GPU
-
Updated
Jul 27, 2026 - Python
Practical implementation and comparison of SFT, Reward Modeling, PPO, DPO, ORPO, and LLM-as-a-Judge for preference alignment with Qwen2.5-0.5B-Instruct.
-
Updated
Aug 18, 2026 - Jupyter Notebook
This is the DPO Group plugin for WooCommerce.
-
Updated
Jun 29, 2021 - PHP
TrainSight: Sub-millisecond pre-flight LLM dataset profiler predicting HBM OOM crashes, padding waste, and attention entropy dispersion before GPU allocation. Built on the Two-Factor Law of LLM Compute.
-
Updated
Aug 4, 2026 - Python
Reproducible QLoRA SFT and DPO experiments for open-weight LLM reasoning, with locked selection, matched controls, and audit-grade provenance.
-
Updated
Aug 7, 2026 - Python
A React-based tool for constructing fine-tuning datasets with list and grid forms, featuring the ability to download and upload data as JSONL files. This project leverages the react-declarative library to create dynamic, interactive forms for defining user inputs, preferred outputs, and non-preferred outputs, along with associated tools
-
Updated
Apr 7, 2025 - TypeScript
Same LoRA DPO recipe across model families (Qwen + Mistral) on UltraFeedback, measured with the correct reference-relative metric. Leads with failure analysis.
-
Updated
Jun 19, 2026 - Python
Fine-tune Llama-3.2-3B (QLoRA/Unsloth) to answer Singapore ElderShield questions closed-book: recall facts, abstain on unknowns, keep general knowledge, stay warm. SFT then DPO on a T4.
-
Updated
Aug 14, 2026 - Jupyter Notebook
[EMNLP 2026] Official code of "Language Chain in Alignment: Cross-lingual Ranking Preference Optimization"
-
Updated
Aug 27, 2026 - Python
Real-world evaluation framework for autonomous pentesting LLMs — 12-dim safety×performance rubric, SFT/DPO/GRPO post-training comparison, MITRE ATT&CK mapping. MCS Capstone @ USyd.
-
Updated
May 23, 2026 - Python
🏟️ Modern RL algorithms from scratch — from Q-Learning to GRPO — with clean PyTorch code and interactive notebooks. Compare PPO vs DPO vs GRPO for LLM alignment.
-
Updated
Mar 30, 2026 - Python
Multi-domain LLM fine-tuning and serving pipeline (LoRA, QLoRA, DPO) — 70% cheaper than proprietary APIs across 5 domain models.
-
Updated
Jul 15, 2026 - Python
Add this topic to your repo
To associate your repository with the dpo topic, visit your repo's landing page and select "manage topics."