Skip to main content

SNAG-Bench

The Quality Certifier. Temporal reasoning benchmark for LLMs. 60 adversarial tasks, 5 scoring axes, 3 difficulty tiers. Designed to stay hard through 2030. Measures Causal Resolution: Coverage x Convergence.

GitHub

timepointai/timepoint-snag-bench — Apache-2.0, Python, Click CLI

Detailed Docs

Full task reference, scoring methodology, and evaluation docs

5 Scoring Axes

Usage

SNAG-Bench is a local CLI tool — not a deployed service.

Task Tiers

Causal Resolution

The composite metric: Coverage x Convergence
  • Coverage — how much of the relevant temporal space is represented
  • Convergence — how consistent the rendered outputs are across runs