What is Vision Language Models?
Vision-language models (VLMs) jointly process images and text for tasks like captioning, visual QA, and grounded instruction following.
Related Giskard articles
Red-team multimodal agents
Image prompts can jailbreak or exfiltrate - include multimodal probes in Giskard suites. See 50+ adversarial probes.
Further reading
Authoritative reference: related vision-language background.