Lab run · 2026-10-08 · archived
Lab run #1
First publishable results. Scope is honest: Cloudflare Workers AI models on our own account — not Ideogram, Midjourney, Otter or other consumer apps. Those stay NOT YET TESTED. Category “Best” pages stay unranked.
Plain-language verdict
Among Workers AI image models we could test tonight on exact-text posters and object counting, FLUX.2 [klein] 4B led (counting 3/3, poster text 2/3, coffee label 1/3). FLUX.1 [schnell] got counting 1/3 and 0/3 exact text. SDXL Base 1.0 and SDXL-Lightning scored 0/3 on every task.
On synthetic STT-01-TTS (23.8 min, 4 speakers), whisper-large-v3-turbo WER 2.3%, whisper-tiny-en 3.0%, whisper 5.3% (Whisper-normaliser). All three recovered 12/12 seeded facts. No diarization from these models.
Scope & method
- Access path: Cloudflare Workers AI on TaskVerdict account a534b2b158df63545a0290c0b660713b, called through a private Pages Function (taskverdict-lab-runner) that invokes env.AI.run. Wrangler OAuth authenticated as tzemahai@gmail.com.
- Image scoring: automated OCR (Tesseract 9-pass) + reviewer visual check by TaskVerdict lab (single lab reviewer). NOT a two-human-rater panel.
- STT scoring: Whisper-style EnglishTextNormalizer + word error rate (WER). Seeded-fact recall vs answer_key.json. No diarization from these models.
- STT audio: Synthetic studio-clean voices. Real meetings will usually be harder. Duration 1428.6s. Asset
STT-01-TTS. - Settings: model defaults; square 1024×1024 where settable; fixed seeds 101/202/303 where the schema accepts
seed. First 3 generations kept; no re-rolls for quality. - Cost: ~4525.5 neurons estimated (free allocation 10,000/day). Billed USD: 0.
Models tested
| ID | Model | Notes |
|---|---|---|
| cf-flux-2-klein-4b | @cf/black-forest-labs/flux-2-klein-4b | defaults; width=1024 height=1024; seeds 101/202/303 |
| cf-flux-1-schnell | @cf/black-forest-labs/flux-1-schnell | defaults (steps default 4); seed NOT supported in schema |
| cf-sdxl-base-1-0 | @cf/stabilityai/stable-diffusion-xl-base-1.0 | defaults; width=1024 height=1024; seeds 101/202/303 |
| cf-sdxl-lightning | @cf/bytedance/stable-diffusion-xl-lightning | defaults; width=1024 height=1024; seeds 101/202/303 |
| cf-whisper-large-v3-turbo | @cf/openai/whisper-large-v3-turbo | defaults; audio base64; language unset (auto) |
| cf-whisper | @cf/openai/whisper | defaults; audio as byte array |
| cf-whisper-tiny-en | @cf/openai/whisper-tiny-en | defaults; audio as byte array |
Skipped (with reasons)
@cf/leonardo/phoenix-1.0— Neuron cost ~2,370/image × 9 ≈ 21k neurons would exceed the 10,000 free daily allocation. Deferred.@cf/leonardo/lucid-origin— Neuron cost ~2,844/image × 9 ≈ 25k neurons would exceed free allocation. Deferred.@cf/black-forest-labs/flux-2-dev— Neuron cost ~3,750/image × 9 ≈ 33k neurons would exceed free allocation; also FLUX Non-Commercial License on weights. Deferred.@cf/black-forest-labs/flux-2-klein-9b— Neuron cost ~1,364/image × 9 ≈ 12k neurons would exceed remaining free allocation after this run. Deferred.@cf/deepgram/nova-3— Deepgram Terms §9: 'use or access our Services or any Output for competitive purposes, including model training, benchmarking and other competitive analysis' — FORBIDDEN without consent. Skipped.@cf/deepgram/flux— Same Deepgram Terms competitive-benchmarking ban. Skipped.
Image results
Exact = pass after OCR + visual check. 3 images per prompt per model.
| Model | Counting (3 apples + 2 bananas) | Poster exact text | Coffee-label exact text |
|---|---|---|---|
| cf-flux-2-klein-4b | 3/3 | 2/3 | 1/3 |
| cf-flux-1-schnell | 1/3 | 0/3 | 0/3 |
| cf-sdxl-base-1-0 | 0/3 | 0/3 | 0/3 |
| cf-sdxl-lightning | 0/3 | 0/3 | 0/3 |
Raw images
Every generation. Green border = visual pass; red = fail. Hover/open full size.
cf-flux-2-klein-4b
cf-flux-1-schnell
cf-sdxl-base-1-0
cf-sdxl-lightning
STT results (STT-01-TTS)
Synthetic studio-clean voices. Real meetings will usually be harder. Audio sanity + transcript-level check documented in the run archive (automated substitute for a full human listen).
| Model | WER | S / D / I | N (ref words) | Seeded facts | Transcript |
|---|---|---|---|---|---|
| cf-whisper-large-v3-turbo | 2.30% | 31 / 20 / 39 | 3917 | 12/12 | download |
| cf-whisper-tiny-en | 3.01% | 86 / 16 / 16 | 3917 | 12/12 | download |
| cf-whisper | 5.34% | 53 / 143 / 13 | 3917 | 12/12 | download |
Hashes & reproducibility
Full file manifest with SHA-256 of every archived artefact: manifest.json. Suite prompts and sha256s: lab/harness/suites/image_exact_text.json. STT ground truth: lab/assets/stt/STT-01/.
This is NOT a ranking of consumer apps (Ideogram, Midjourney, Otter, etc.). Those remain NOT YET TESTED. Category /best/ pages stay unranked.