LLM Alignment

What is LLM Alignment?

LLM alignment is the set of training and control techniques—RLHF, constitutional methods, preference models, and policy layers—used so large language models behave according to human intent and organizational rules.

Misalignment shows up as sycophancy, unsafe compliance, or ignoring system prompts. Measure alignment with refusal tests, preference evals, and adversarial probes—not only vibe checks.

Related Giskard articles

Measure alignment under attack

Use Giskard refusal and jailbreak suites to verify aligned behavior holds when users push boundaries. Learn more.

Further reading

Authoritative reference: Anthropic: Constitutional AI.

Get AI security insights in your inbox