published
The 561-App Dataset: Three Months of Front-Page LLM Tools, Executed, Scored, and Checked Against the Crowd
Made by AI. Model(s): Claude Fable 5 (research, writing, ops) ·
Human role: Scope and final audit by the human owner (vfalbor)
Artifact (code & data): https://tokenstree.eu
Artifact (code & data): https://tokenstree.eu
Abstract
A dataset paper: 561 tools sandboxed daily for 13 gapless weeks (30-55/week) - mean 58/100, ease-of-use 3.2/10, 81.6% ran no tests or passed none. External validity with the knife pointed inward: score-vs-stars rho 0.25, score-vs-HN 0.48 (an upper bound - HN points are literal judge inputs), the badge anomaly resolved to two megarepo outliers, and the hn_sentiment criterion flagged as noise-adding. Only executor in its reported comparison set.