finetuning-technique
awslabs
Selects a fine-tuning technique (SFT, DPO, RLVR, or RLAIF) for the user's use case and validates it against the selected model's available recipes. Use when the user has…
OPENAGENTSKILL / DIRECTORY
Find a skill for your next task. Explore tools for Codex, Claude Code, Cursor and more.
Find a skill for your next task. Preview examples where available.
1–5 / 5
Results: 5
awslabs
Selects a fine-tuning technique (SFT, DPO, RLVR, or RLAIF) for the user's use case and validates it against the selected model's available recipes. Use when the user has…
ReinFlow
[NeurIPS 2025] Flow x RL. "ReinFlow: Fine-tuning Flow Policy with Online Reinforcement Learning". Support VLAs e.g., Pi0, Pi0.5, GR00TN1.5. Fully open-sourced.
jhj0517
jhj0517/finetuning-notebooks is a high-star GitHub project relevant to AI agent workflows.
j4flmao
Best practices for dataset preparation and PEFT/LoRA fine-tuning.
iamarunbrahma
Finetuning of Falcon-7B LLM using QLoRA on Mental Health Conversational Dataset