Agent skill

nlp-alignment

Best practices for LLM alignment techniques including RLHF, DPO, and instruction tuning. Use when working on alignment or safety.

Stars 11,027
Forks 1,262

Install this agent skill to your Project

npx add-skill https://github.com/aiming-lab/AutoResearchClaw/tree/main/researchclaw/skills/builtin/domain/nlp-alignment

Metadata

Additional technical details for this skill

author
researchclaw
version
1.0
category
domain
priority
4
references
Ouyang et al., Training language models to follow instructions, NeurIPS 2022; Rafailov et al., DPO, NeurIPS 2023
trigger keywords
alignment,rlhf,dpo,reward model,preference,instruction tuning,safety
applicable stages
9,10

SKILL.md

LLM Alignment Best Practice

Methods:

  • RLHF: Train reward model → PPO fine-tuning (complex but powerful)
  • DPO: Direct preference optimization (simpler, no reward model needed)
  • GRPO: Group relative policy optimization
  • SFT: Supervised fine-tuning as alignment baseline

Training recipe:

  • Start with SFT on high-quality instruction data
  • DPO: lr=5e-7, beta=0.1, batch_size=64
  • PPO: lr=1e-6, clip=0.2, KL coeff=0.02
  • Use reference model for KL penalty
  • Evaluate on safety benchmarks (TruthfulQA, BBQ, etc.)

Common pitfalls:

  • Reward hacking: model finds shortcuts to high reward
  • Mode collapse: model generates repetitive outputs
  • Catastrophic forgetting: loses general capabilities

Expand your agent's capabilities with these related and highly-rated skills.

aiming-lab/AutoResearchClaw

scientific-visualization

Publication-ready scientific figure design with matplotlib and seaborn. Use when creating journal submission figures with proper formatting, accessibility, and statistical annotations.

11,027 1,262
Explore
aiming-lab/AutoResearchClaw

hypothesis-formulation

Structured scientific hypothesis generation from observations. Use when formulating testable hypotheses, competing explanations, or experimental predictions.

11,027 1,262
Explore
aiming-lab/AutoResearchClaw

scientific-writing

Academic manuscript writing with IMRAD structure, citation formatting, and reporting guidelines. Use when drafting or revising research papers.

11,027 1,262
Explore
aiming-lab/AutoResearchClaw

a-evolve

Apply A-Evolve's agentic evolution methodology to improve AI agent performance across runs. Use when the user wants to diagnose agent failures, generate targeted skills from error patterns, evolve system prompts, or accumulate episodic knowledge. Works standalone or inside AutoResearchClaw pipelines. Triggers on: "evolve", "self-improve", "diagnose failures", "generate skills from errors", "what went wrong and how to fix it", or any mention of A-Evolve.

11,027 1,262
Explore
aiming-lab/AutoResearchClaw

chemistry-rdkit

Computational chemistry with RDKit for molecular analysis, descriptors, fingerprints, and substructure search. Use when working with SMILES, drug discovery, or cheminformatics tasks.

11,027 1,262
Explore
aiming-lab/AutoResearchClaw

literature-search

Systematic literature review methodology including search strategy, screening, and synthesis. Use when conducting literature reviews or writing background sections.

11,027 1,262
Explore

Didn't find tool you were looking for?

Be as detailed as possible for better results