Giskard · freemium
Giskard
model-eval
What it does
Open-source testing framework for ML and LLM applications. Generates test suites that probe for performance, robustness, hallucination, bias, and prompt injection; supports a hub for collaboration and result tracking.
What it doesn't do
The auto-generated tests can be a starting point but typically need human curation for domain-specific assessments. Some of the deeper enterprise features (hub, collaboration) are paid.
Best for
- ML engineering teams adding LLM/agentic testing to existing ML test discipline
- First-pass risk assessment of a new LLM application
Not for
- Replacing a dedicated AI red team — augments, does not substitute
Pricing
Free OSS; paid for the hosted hub.
Integration effort
medium
Team skill required
Python; existing ML testing experience helpful.
Last reviewed: 2026-05-01