SNAG-Bench
The Quality Certifier. Temporal reasoning benchmark for LLMs. 60 adversarial tasks, 5 scoring axes, 3 difficulty tiers. Designed to stay hard through 2030. Measures Causal Resolution: Coverage x Convergence.GitHub
timepointai/timepoint-snag-bench — Apache-2.0, Python, Click CLIDetailed Docs
Full task reference, scoring methodology, and evaluation docs
5 Scoring Axes
Usage
SNAG-Bench is a local CLI tool — not a deployed service.Task Tiers
Causal Resolution
The composite metric: Coverage x Convergence- Coverage — how much of the relevant temporal space is represented
- Convergence — how consistent the rendered outputs are across runs