What is LLM Alignment?
LLM alignment is the set of training and control techniques—RLHF, constitutional methods, preference models, and policy layers—used so large language models behave according to human intent and organizational rules.
Misalignment shows up as sycophancy, unsafe compliance, or ignoring system prompts. Measure alignment with refusal tests, preference evals, and adversarial probes—not only vibe checks.
Related Giskard articles
Measure alignment under attack
Use Giskard refusal and jailbreak suites to verify aligned behavior holds when users push boundaries. Learn more.
Further reading
Authoritative reference: Anthropic: Constitutional AI.