NVIDIA Garak ships 120+ vulnerability categories, the broadest open-source probe library for foundation-model testing. That breadth is also why agent teams outgrow it: tool-calling misuse, MCP abuse, and multi-turn jailbreaks sit outside Garak's model-level, single-turn design.
What is NVIDIA Garak?
Garak is an open-source LLM vulnerability scanner from NVIDIA with a research pedigree. It is model-agnostic, supporting Hugging Face, OpenAI, Bedrock, REST endpoints, and GGUF, and ships the broadest probe library in our 2026 agent red teaming tools guide. Teams use it for systematic model-level red teaming without a commercial license.
What our guides highlight
- 120+ vulnerability categories, the broadest probe library in our comparison
- Model-agnostic coverage across Hugging Face, OpenAI, Bedrock, REST endpoints, and GGUF
- Strong research pedigree from NVIDIA's security research community
- Fully open-source with no commercial license required
Why teams look elsewhere
Garak earns its place as the open-source baseline for foundation-model scanning. With 120+ vulnerability categories, it offers the broadest probe library in our guides. Model-agnostic support across Hugging Face, OpenAI, Bedrock, REST endpoints, and GGUF fits heterogeneous inference stacks. Strong research pedigree and a fully open-source license make it a natural choice for research teams and budget-conscious engineering orgs.
Teams widen the search when risk shifts from models to agents. Garak is designed for model testing, not agent testing, with no tool-calling or MCP support. Attacks are primarily single-turn and static, and automated attack generation (atkgen) remains largely prototype and stateless. There are no collaboration or workflow features, full scans can be resource-intensive, and there is no vulnerability-to-fix pipeline.
How we evaluate tools
We apply the framework from our 2026 agent red teaming guide and 2025 comparison: security and quality tested together, agent and multi-turn coverage beyond model-only scans, and a path from findings to tasks, regression tests, and guardrails. We also weigh whether product and domain teams can participate, not only a security-led scanner. The notes below come from those guides and our matrix.
Why teams choose Giskard
Giskard is the stronger fit when production risk lives in agent behavior, not model baselines alone. Giskard extends red teaming to agent tool calls, interaction history, and adaptive multi-turn attacks such as GOAT, threats Garak's probe library was not designed to exercise. Findings flow into tasks, regression tests, and guardrails, and Giskard Hub supports cross-functional collaboration beyond a CLI scanner. Keep Garak for model baselines; choose Giskard when agent tool misuse across conversations is the core threat.
Other alternatives
Promptfoo
PR-native YAML configs, CI/CD integration, and early MCP vulnerability plugins. Strong when engineers want declarative eval in pull requests. See Promptfoo alternatives.
Microsoft PyRIT
Composable orchestrators from Microsoft's AI Red Team with native Azure integration. Fits security teams building custom attack pipelines in Python. See PyRIT alternatives.
Confident AI DeepTeam
Python library with 40+ vulnerabilities, attack methods, and built-in guardrails plus an agentic red-teaming module. See DeepTeam alternatives.
Lasso Security
Agent inventory, MCP scanning, and pre-attack reconnaissance with a large enterprise attack library. See Lasso Security alternatives.
When NVIDIA Garak still makes sense
- You need the broadest probe library (120+ vulnerability categories) for model-level testing
- Model-agnostic coverage across Hugging Face, OpenAI, Bedrock, REST, and GGUF endpoints matters
- Open-source, research-grade tooling with NVIDIA's research pedigree fits your stack and budget
- Foundation-model baselines must run without a commercial license or vendor lock-in
When Giskard is the better fit
- Your application uses tool calling, MCP servers, or multi-step agent workflows
- Adaptive multi-turn attacks (GOAT) better match your threat model than static probes
- Non-technical stakeholders must participate in scenario design and review
- You need findings to become tasks, regression tests, and guardrails, not CSV output
- Full Garak scans are too resource-intensive for your CI environment
Bottom line
Giskard is the stronger fit when agent tool calls, multi-turn jailbreaks, and a fix pipeline matter more than probe breadth alone. Garak still wins when open-source, model-agnostic foundation-model scanning with 120+ categories is the primary job.
