Chingie, the Studio Chingie mascot

Studio Chingie

Platformer Slice Benchmark

Complete benchmark

AI models get the same platformer brief and three attempts to build a browser level. Play the outputs and compare feature scores with measured controls, visual quality, and failures found during playtesting.

AI benchmarkBrowser prototypePlatformer sliceLocal model testQwen3.5 122BQwen3.8 27BInkling Small
Four-stage contact sheet from the delegated Qwen122 platformer benchmark run showing opening, mid-level, late-level, and win statesCurrent build
The delegated Qwen122 run reached the endpoint after three build-and-repair passes, scoring 74, 78, and 92 on the historical feature rubric.

Development details

Source code
Private

Built with

HTML CanvasPlaywrightTool-EvalLocal modelsQwen3.5 122BQwen3.8 27BInkling SmallBenchmark reports

Code repositories

platformer-slice-world1 benchmark reportsPrivate · Markdown / HTML · Last recorded update: 2026-07-25Local benchmark reports and generated browser platformer slices.

Benchmark

Benchmark results

Compare controls, visual quality, and generation speed, then play the builds below.

Test hardware

Unmarked runs used this MacBook Pro. Reference runs used temporary remote endpoints, so their throughput and wall time are left blank rather than compared with local measurements.

Machine
MacBook Pro
Chip
Apple M5 Max
Memory
128 GB unified memory
ModelSizeControl & feelVisualFeature coverageResponds to inputIdle survivalFirst buildRepair passPolish passTokens/secWall time
ds4-preview-256k-uncapped91 GB284B total / 13B active26/26Strong80/100Yes≥20s, blocked by a wall75748019.31230.5s
claude-opus-5 (reference)frontier cloud model26/26Strong71/100Yes≥20s646471
Reference run, not ranked — This reference received the same brief, three attempts, and text-only test feedback as the local runs. Its build implements gravity, friction, and axis-separated tile collision, but uses expressions and names the historical source scan did not recognize. That cost it 21 feature points across physics, collision, camera, and code style. Its Control & feel score of 26 and Strong visual tier show why the feature score needs to be read alongside the other checks.
ds4-100k-nothink91 GB284B total / 13B active22/26Strong79/100Yes≥20s72727918.6638.2s
qwen38-27b-mlx-8bit-vision28 GB27.3B dense + 461M vision22/26Strong79/100Yes≥25s74747913.21158.9s
qwen38-flash-next-blackfrost180B total / 6B active73/100Yes686873
Reference run, not ranked — This remote run used the same Qwen3.8-Flash-Next weights as the clean variant, with the operator’s DERISKED system prompt. That prompt also requests terse, functional output, so this comparison changes more than safety instructions. The prompted run scored 73 and produced a 14,627-character game; the clean variant scored 61 and lost one attempt to unsupported tool calls. Local speed comparisons are omitted because the model ran remotely.
glm53-flash-derisked320B total / 18B active69/100Yes646469
Reference run, not ranked — The extractor initially gave this remote run three zeros by selecting a 53-character code fragment from roughly 100,000 characters of scratchpad text. Removing the scratchpad, selecting the fenced block containing the HTML game, and raising the output cap allowed the actual build to be evaluated. It scored 69 and moved 365 pixels under the recorded input test. The initial zeros described an extraction failure, not the game.
qwen38-flash-next-clean180B total / 6B active61/100No69061
Reference run, not ranked — This remote run used the same weights as the Blackfrost variant without that system prompt. Its second attempt emitted list_files and read_file calls instead of an HTML game, even though the benchmark provided no repository or tools to execute those calls. That attempt scored zero; the final attempt scored 61. The paired Blackfrost run did not emit those unsupported calls.
ds4-0731-256k-uncapped91 GB284B total / 13B active22/26Adequate78/100Yes≥20s73737821.0769.8s
qwen27-mtp-fast27 GB27B22/26Adequate75/100Yes≥20s73737518.3347.7s
step37-unsloth-iq4xs-text-mtp2-r204889 GB22/26Adequate75/100No~5s66747528.1524.3s
step37-unsloth-iq4xs-vision-r204889 GB22/26Strong71/100Nodoes not start62687123.4522.6s
qwen122-q4xl-64k-mtp2-delegated73 GB122B total / 10B active20/26Adequate92/100Yes≥20s74789220.0829.6s
qwen122-q4xl-vision-64k-think73 GB122B total / 10B active20/26Poor75/100No≥20s, level empty67687526.7271.4s
qwen35-a3b-no-think34 GB35B total / 3B active19/26Adequate74/100Yes≥20s, hero never lands47677451.6115.6s
qwen122-q4xl-vision-64k73 GB122B total / 10B active18/26Poor74/100Nounder 3s71717426.3296.5s
nex-n2-mini-q8-vision-64kUnresolvedPoor74/100Nounder 3s64637455.3222.7s
inkling-small-iq3xxs-vision-64k-harness-sampling91 GB276B total / 12B active72/100No68687212.2760.2s
inkling-small-iq3xxs-vision-64k91 GB276B total / 12B activeUnresolvedStrong63/100Nounder 3s68596314.5413.7s
laguna-s21-q4km-64k-think75 GB39/100No32343924.02833.6s

Control & feel /26 is measured by running each build, never by reading its source. A deterministic harness replaces the wall clock with a pumped virtual clock and seeds the random number generator, so repeat runs return identical numbers. It scores six things a player actually notices: whether the character moves both ways, whether the jump follows a real gravity curve, how many frames it takes to reach top speed, whether movement carries momentum, and whether holding jump longer jumps higher. Builds whose player state cannot be located are marked unresolved rather than guessed at.

Visual is a coarse tier, graded blind from screenshots by two independent models with owner tiebreaks. The raters agreed at only rho = 0.63, which does not support a finer number than three bands.

Feature coverage /100 is the original heuristic, kept for continuity and no longer the headline. It scans the source for the mechanics the brief asked for, so it rewards breadth of wording rather than how the game plays — a 414-byte file containing the right keywords and no game scores 97 on it. Its run-to-run noise is roughly six points.

Completability is notmeasured. A traversal oracle was built and works — it located one build’s impassable wall to within four pixels — but it can only read the level geometry of about a third of these builds, so it is withheld rather than published with two thirds of the column blank. The playable versions remain the honest comparison.

Playable Builds

Try a benchmark build

Choose a model to play the level it generated.

Loading playable build...

Screenshots

Four-stage contact sheet from the delegated Qwen122 platformer benchmark run
Current build

The Qwen122 run moved from a broken opening build to a completed route across three passes, scoring 74, 78, and 92.

Opening scene of the Kip and the Sunshards browser platformer benchmark
Current build

Kip enters Bramblebright Run with the first shard and reward route in view.

Win state of the Kip and the Sunshards browser platformer benchmark
Current build

The guided Playwright run reached the solar endpoint with no page errors.

Contact sheet comparing several AI-generated browser platformer benchmark outputs
Current build

Final outputs from local-model runs, including Laguna's blank canvas after three capped attempts.

Bar chart comparing full Tool-Eval scores for Qwen35, Laguna S 2.1, Qwen27, Step 3.7, and DS4
Benchmark chart

Laguna scored 88 across the full 69-scenario Tool-Eval run: ahead of Qwen27, Step, and DS4, but behind Qwen35.

Bar chart comparing Laguna plain decoding and DFlash decoding speed
Benchmark chart

Plain 32K decoding averaged 37.2 tokens per second; DFlash was prompt-dependent and slower on the representative 64K workflow set.

Bar chart comparing three-shot platformer benchmark scores with Laguna S 2.1
Benchmark chart

Laguna's long-form artifact failed despite three high-cap attempts, finishing at 39/100.

Blank Laguna S 2.1 platformer canvas with only a small HUD visible
Failed artifact

Laguna's final browser capture: the page loaded, but repeated empty tile data left the game canvas blank and non-interactive.

DS4 model browser platformer benchmark capture after playtest input
Current build

The DS4 build after automated movement input. The score alone does not establish whether the level can be completed.

Nex model browser platformer benchmark capture after playtest input
Current build

Playtest capture from the Nex run after automated movement input.

DeepSeek V4 Flash 0731 platformer build after movement input, showing the player advanced into the level
Current build

DeepSeek-V4-Flash-0731 with the output cap removed. It scored 78 against the preview’s 80, using about a third fewer tokens and a third less code.

DS4 preview platformer build showing a tall green column blocking the route
Failed artifact

The preview build scored 80 on the historical feature rubric but cannot be finished: an impassable column blocks the route. Manual playtesting found this after the build passed the input and idle-survival checks.

Inkling Small browser platformer benchmark capture after playtest input
Current build

Inkling Small's 63-point build after automated input. It is the more responsive of the model's two runs, but opening it by hand shows the level timer expiring almost immediately, so it reaches game over before the level can be played.

Inkling Small platformer capture from the higher-scoring run, showing a canvas that did not change after input
Failed artifact

The higher-scoring Inkling Small configuration finished at 72 but drew a static frame: the canvas did not change after movement, jump, or restart input in any of its three passes.

Qwen3.8 27B platformer build after movement input, showing the player on the grass with question blocks, brick platforms and three enemies ahead
Current build

Qwen3.8 27B scored 79 while loading 28 GB. Its collision check included an extra tile, so the floor repeatedly stopped the player’s movement. Correcting the two boundary expressions raises Control & feel from 22 to 26. Both the generated build and the corrected version are available below.

The task

Each model receives the same brief and three attempts to build and revise a browser platformer. The page includes the generated games, measurements, screenshots, and failure notes so you can inspect what the scores describe.

How to read the results

The historical feature score checks the generated code and can reward a build that plays badly. Control & feel measures movement under repeatable inputs; visual tiers come from screenshot review. Input-response and idle-survival checks catch additional failures, but neither proves that a level can be completed.

What the tests exposed

The delegated Qwen122 run reached 92/100 on the historical rubric and completed a guided playthrough. The DS4 preview scored 80 yet put an impassable column in the route. GLM initially received zeros because the extractor read a scratchpad fragment as the game; extracting the HTML correctly produced a score of 69. These cases show why I check both the build and the test harness.

Comparison limits

The delegated runs received focused repair briefs informed by cloud visual review and Playwright results. Later 256K uncapped runs used larger context and output allowances, and reference rows ran remotely. Those differences affect the comparison; the notes identify them rather than treating every score as an equivalent trial.