[NeurIPS 2024] SimPO: Simple Preference Optimization with a Reference-Free Reward
-
Updated
Feb 16, 2025 - Python
[NeurIPS 2024] SimPO: Simple Preference Optimization with a Reference-Free Reward
[Paper][ACL 2024 Findings] Knowledgeable Preference Alignment for LLMs in Domain-specific Question Answering
Video Generation Benchmark
Code for "ReSpace: Text-Driven Autoregressive 3D Indoor Scene Synthesis and Editing"
DPO-Shift: Shifting the Distribution of Direct Preference Optimization
[NeurIPS 2024] Official code of $\beta$-DPO: Direct Preference Optimization with Dynamic $\beta$
Source code for "A Dense Reward View on Aligning Text-to-Image Diffusion with Preference" (ICML'24).
[ICCV 2025] Official repository of "Mitigating Object Hallucinations via Sentence-Level Early Intervention".
[ICML 25] "Preference Optimization for Combinatorial Optimization Problems"
[ICLR 2025] Official code of "Towards Robust Alignment of Language Models: Distributionally Robustifying Direct Preference Optimization"
[ICLR 2026] Official repository of "Uni-DPO: A Unified Paradigm for Dynamic Preference Optimization of LLMs".
[ICLR 2025] Bridging and Modeling Correlations in Pairwise Data for Direct Preference Optimization
[ICML 2025] TGDPO: Harnessing Token-Level Reward Guidance for Enhancing Direct Preference Optimization
CapField-OPD: Learning Continuous Capability Fields via Joint-Anchored Multi-Teacher On-Policy Distillation for Flow Models
[NeurIPS 2025] Ranking-based Preference Optimization for Diffusion Models from Implicit User Feedback
A curated collection of papers & repos on Auto-Skill (self-evolving agents) and Auto-Rubric (rubric learning from preferences) for LLM alignment & customization.
[ECCV2026] RefReward-SR: LR-Conditioned Reward Modeling for Preference-Aligned Super-Resolution
LLM-Driven Preference Data Synthesis for Proactive Prediction of the User’s Next Utterance in Human–Machine Dialogue
Production-grade SFT + ORPO fine-tuning scripts for small reasoning & instruction models. Built with Unsloth + TRL, stabilized for single-GPU (RTX 4090) training with long thinking tokens. More models coming.
A living, evidence-based tracker of personalized LLM & agent research — benchmarks, methods, and surveys across preference alignment, long-term memory, multimodal, tool use, safety and privacy.
To associate your repository with the preference-alignment topic, visit your repo's landing page and select "manage topics."