Quality
👍/👎 rate on finding cards, sourced from finding_rated events. Per metrics-strategy §7 the operational target for the thumbs-down rate is <5%; sustained drift above that is the model-degradation alarm. The divergence strip below flags buckets where users and the judge disagree — click a bucket to filter. Reason text lives in Firestore — drill into a session for the qualitative read.
Loading…
Trend
(?)👍 rate by prompt version
Judge verdicts
Divergence (users vs. judge) — click a bucket to filter
Insights (last 7 days)
(?)
Click "Generate insights" to synthesize the last 7 days of feedback.