The most interesting frontier-model showcase is not another benchmark chart. It is a project with enough state to surprise you, enough control to let you push back, and enough public evidence to distinguish a build from a launch claim. On that standard, MiniTown is the best GPT‑5.6 Sol example and VibeCAD is Claude Fable 5's boldest exhibit.
The ranking rule
Each list uses the same four editorial tests: interaction depth—does the artifact expose a world rather than a screenshot; product coherence—do its parts reinforce one clear experience; model leverage—does the work reveal something beyond generic code generation; and public evidence—can an outsider inspect, run, or at least trace the claim to its first-party source? Rank is a recommendation for builders seeking inspiration, not a measured model evaluation.
Interaction depth
A system with state, controls, and consequences beats a static composition.
Product coherence
One memorable job and a complete loop beat a bag of disconnected features.
Model leverage
The build should expose planning, taste, persistence, or verification—not just volume.
Public evidence
Live artifacts, prompts, source, and run receipts increase confidence in that order.
The top five GPT‑5.6 Sol projects
All five are first-party showcase entries described as built with Codex and GPT‑5.6. Their pages expose the brief; each also links to a live build. That does not make them independent evaluations, but it does make them unusually useful design references.
Best overall · living systems
MiniTown
A cozy miniature city where zoning creates roads, buildings pass through visible construction stages, residents acquire homes and destinations, traffic moves, and the light shifts through a day. Its charm comes from simulation state, not a hero animation: you place a few causes and watch a legible town emerge.
- Why it ranks first
- The strongest combination of systemic depth, readable causality, and emotional product taste.
- Evidence
- Official project page, build prompt, credited builder, and linked live app. No public source repository on the page.
Best real tool · 3D design
Material Lab
A browser studio for tuning glass, metal, ceramic, and fabric across color, roughness, translucency, lighting, environments, camera angles, comparisons, and saved presets. Unlike a glossy landing page, every visual decision is connected to a control a designer can manipulate immediately.
- Why it ranks here
- It feels closest to a useful creative instrument: focused, reversible, and built around fast visual feedback.
- Evidence
- Official project page, detailed prompt, and live app. The page does not expose the implementation repository.
Best polish loop · multi-agent build
Tiny Rails Rollercoaster
A miniature 3D coaster simulator with eight route combinations and four driving temperaments, from assisted to deliberately chaotic. The published iteration log is the real showcase: geometry, banking, acceleration, recovery, environment, camera behavior, and performance were refined as distinct product seams.
- Why it ranks here
- It demonstrates that coordinated iteration and rendered-result inspection matter more than a heroic first prompt.
- Evidence
- Official project page, live app, prompt, and a five-stage iteration summary; no public source repository.
Best toy · physics and restraint
Glass Towers
A minimalist balancing game where translucent pieces with different masses and centers of gravity drop onto a narrow pedestal. The rule is grasped instantly, but shape, momentum, lighting, impact, score, preview, and failure state give the toy enough texture to sustain attention.
- Why it ranks here
- The concept is unusually clean: one verb, one risk, and a visual material system that supports the mechanic.
- Evidence
- Official project page, full initial prompt, and linked live app. Source and run history are not public there.
Best commerce interaction · accessible delight
Field Day
A picnic storefront whose signature basket can accept dragged products or an accessible Add action, then unfold into a tidy gift-set review before entering a persistent local cart. It keeps conventional browsing legible while reserving playfulness for one interaction the customer can understand.
- Why it ranks here
- It treats touch, keyboard, responsive verification, and commerce clarity as part of the visual idea.
- Evidence
- Official project page, live app, detailed brief, and explicit desktop/mobile verification request; no public code link.
The top five Claude Fable 5 projects
These are best read as ambitious exhibits, not reproducible case studies. Anthropic's launch page names and describes each demonstration, but it does not provide a separate public project page, repository, prompt, or full trace for the five below. The ranking therefore emphasizes what each experiment teaches, while grading every evidence packet as provider-reported.
Boldest recursive tool · CAD
VibeCAD + printable model
Fable 5 reportedly created a browser-based CAD editor, including an AI copilot, then used that environment to design a complete 3D-printable object. The recursion is the point: build the tool, inhabit the tool, and produce a constrained artifact whose geometry must survive outside the chat window.
- Why it ranks first
- It crosses from interface generation into toolmaking and manufacturable geometry—the strongest project-shaped claim in the Fable set.
- Evidence
- Anthropic launch description and embedded demonstration only; no standalone app, source, print file, or independent reproduction linked.
Best scientific explainer · first principles
Solar eclipses
A solar-system simulation whose orbital motion was derived from physics rather than keyframed for effect, then used to predict eclipses. It is the right shape for an AI explainer: a visual surface attached to a model of the world, with a future event that can in principle falsify the implementation.
- Why it ranks here
- The project ties code, visual explanation, mathematical structure, and an externally checkable phenomenon together.
- Evidence
- Anthropic launch description and embedded visual only; equations, validation results, code, and a live simulator are not linked.
Best long-horizon agent · world state
Autonomous Factorio factory
The model plays Factorio by planning production, routing resources, constructing an automated factory, and recovering inside a world where early errors compound. Factorio is an unusually honest agent environment because progress lives in the factory state, not in the model's narration of its own success.
- Why it ranks here
- It tests persistence, spatial reasoning, resource planning, and error recovery in one legible system.
- Evidence
- Anthropic launch description and embedded demonstration; the harness, save file, action trace, and full run are not linked.
Best perception loop · vision only
Pokémon FireRed completion
Anthropic reports that Fable 5 completed Pokémon FireRed from raw screenshots without maps, navigation aids, or privileged game-state tools. The achievement is less about nostalgia than interface generality: perceive pixels, remember a long route, choose an action, and survive thousands of imperfect steps.
- Why it ranks here
- It strips away helper scaffolding and makes the perception-action-memory loop the entire experiment.
- Evidence
- Anthropic description and time-lapse; the complete trajectory, failure count, controls, and reproducible harness are not linked.
Best strange prototype · coded media
Fluid with Classical EDM
A coded fluid simulation synchronized to a classical-EDM remix that Fable 5 also produced in code. The intriguing move is not “AI made music.” It is one generated timing system driving another generated visual system, producing a composition whose behavior can be inspected at the code and rendered-output layers.
- Why it ranks here
- It is the most category-breaking experiment: simulation, rhythm, visual art, and program synthesis in one artifact.
- Evidence
- Anthropic launch description and embedded demo only; audio code, fluid implementation, prompts, and live artifact are not linked.
The useful difference is not “which model wins?”
Products with public handles.
The standout Sol examples expose a live URL, a bounded interaction, and a brief. Their value is immediate inspectability, even when source code is absent.
Experiments with longer arcs.
The Fable examples reach for tool creation, world simulation, and extended autonomous play. Their ambition is high; the public evidence surface is thinner.
Let state carry the story.
Towns, materials, coasters, factories, planets, and games keep the proof outside the model's prose. The artifact changes and the user can see why.
Do not confuse a provider reel with reproduction.
A first-party showcase is a primary source for what the provider claims. It is not independent verification, a full run trace, or a public code audit.
What builders should cook next
- Borrow MiniTown's causal readability: make every user input produce a visible system consequence.
- Borrow Material Lab's reversibility: expose controls, comparison, reset, and saved state before adding more generation.
- Borrow Tiny Rails' iteration log: treat physics, camera, recovery, environment, and performance as separate polish passes.
- Borrow VibeCAD's recursion: let the model create a constrained tool, then judge it by the artifact the tool produces.
- Borrow Factorio's truth surface: prefer environments where progress is externally observable and failure compounds honestly.
- Improve on both providers: publish the prompt, source, trace, live artifact, limitations, and one failed run beside the highlight reel.
Now playable in Chopshopr Labs
We cooked the patterns into one living system.
Causal Forge combines MiniTown-style visible consequences, VibeCAD-style artifact export, Factorio-style autonomous flow, and explicit failure rehearsal. Describe a mission, forge five bounded agents with GPT‑5.6, inject a fault, inspect recovery, then keep the blueprint and rendered world.
Methodology · snapshot July 19, 2026
What was inspected, and what was not inferred.
Source boundary
The Sol shortlist came from OpenAI's official developer showcase and GPT‑5.6 launch materials. We opened the individual project pages and used their descriptions, prompts, iteration notes, credits, and live links. The Fable shortlist came from Anthropic's official Fable 5 launch and model pages. We treated provider descriptions as evidence of provider claims, not as independent validation.
Selection boundary
- Included: named artifacts that reveal a distinctive product, system, or extended agent loop.
- Excluded: benchmark scores, customer testimonials without an inspectable artifact, Mythos-only biology and cyber work, and generic capability claims.
- Not inferred: the exact share of code written by a model, unreported human intervention, production readiness, scientific correctness, or reproducibility beyond the published evidence.
Primary references
- OpenAI GPT‑5.6 launch
- OpenAI developer showcase
- MiniTown project page
- Material Lab project page
- Tiny Rails project page
- Glass Towers project page
- Field Day project page
- Anthropic Fable 5 launch
- Claude Fable model page
Read the companion field notes
This project showcase complements our ranked map of the libraries beneath Fable and Sol demos and the separate OpenAI Build Week contender audit. One piece shows the artifacts, one maps the enabling primitives, and one studies the competition field. None substitutes for a reproducible source release.