A Hundred Playable Demos and Nothing to Ship
AppMagic's H1 2026 casual report carries two numbers worth holding side by side. Developers shipped more than 2,000 Block Puzzle games in six months, up 89% year over year. In the same window, 120 new Match-3 titles launched and exactly one crossed $100K in monthly revenue: a Vietnamese version of Gardenscapes. The only new winner was an old one.
Hybridcasual is what a market looks like after production costs collapse: not more winners, the same handful standing in a much taller pile. AI prototyping now brings that economics to every genre, including the ones where a vertical slice cost six figures and half a year.
The slice got cheap. The decision didn't.
What a demo used to prove
Greenlight processes rest on an assumption nobody wrote down: the prototype itself was evidence. A compelling slice proved craft and commitment. That assumption just died.
A demo anyone can generate in a day proves nothing about the team and nothing about the market. Most gates still weigh the software like currency.
I've sat in the room. A slate review with eight candidates is hard. With thirty, the meeting stops working: attention fragments, the loudest champion wins, and the projects that die are the ones whose sponsors were out that week.
Sure, thirty shots beat eight in a hit-driven business, and AI even puts all thirty in front of real players instead of a conference room. That works exactly as long as someone owns the verdicts.
Optionality without discrimination isn't a portfolio. It's inventory. Anyone can run the test now; reading the results is still the part that takes a career.
The evidence problem got worse, not better
The classic answer is demand-signal testing, old discipline. On the Hungry Shark sequel Primal, we ran a fake store page with real ads and no game behind it: nine Facebook tests across five themes and four art styles, static icons. Statistically significant winners on both questions before a dollar of production moved. Standard F2P practice a decade ago.
The 2026 problem is uglier. The tools flooding your greenlight with demos are flooding the market with evidence: pristine trailers, polished store pages, CPI tests that read as real. Painted-door tests like the Primal run worked because faking the door cost something. Now the doors are free.
Deciding which numbers still mean something is a judgment call, and the people who make that call well have paid for a bad read. The signals that still discriminate come in classes, not KPIs, because any named number gets farmed once it's fashionable:
1. Time. A day-seven return. An unprompted second session.
2. Money. A deposit. A pre-order at real price.
3. Reputation. A creator staking their channel on your unreleased game.
Rotate what you trust; a known metric is a target. Re-pick the signals that count at every kill review instead of writing them once and defending them. AI can manufacture impressions at any volume. It cannot make ten thousand strangers come back Thursday.
Gates are political technology
Projects don't slip through gates because the checklist was vague. They slip through because the CEO likes the IP. Portfolio discipline breaks on behavior, not math.
A gate does its real work before anyone falls in love. Set the evidence bar and name the kill owner while everyone can still think, and the kill fails a standard, not a person. Then put the displaced team on live projects within the month, because a studio where dead projects damage careers stops generating honest pitches within two cycles.
A gate only binds people who agree to be gated. If the CEO's pet project won't face the same bar, stop building gates. The organization already knows the truth.
One kill taught me the limit in person. A new-IP bid from a young team hungry for a first win: charismatic director, goodwill to burn, a game generic at best underneath. Playtests middling, positioning unfindable, peer reviews mixed to poor, and the low-cost team didn't make the opportunity cost low.
Two try-agains postponed the pain. When the evidence finally forced the question, I led the kill in a contentious room on a Friday morning. By Monday morning the game was alive again as if nothing had happened.
No process document showed what actually happened: the studio head had a line to the CEO that no org chart records, and it outranked every gate in the building. My wrong read wasn't the game. It was believing a kill is a decision. A kill is an enforcement, and I hadn't secured it.
The game failed soft launch four months later. The lessons-learned deck came two months after that, and said what the gate had said six months and one try-again earlier. The director had checked out before the kill, a console veteran out of his depth with no way off.
If you're standing inside the project, the tell: a resurrection with no new evidence isn't belief. It means the project is useful to someone above you, as cover, as a favor owed, or to save someone's face. You're not being backed. You're being spent.
That's the other way a studio loses judgment. The first way was the layoffs: two years of cutting the people who carried it while budgeting billions for AI infrastructure. The second is quieter: teaching the people still in the building not to spend it.
Every yes names its no
One criterion straight off our gate sheet, free: once a candidate clears the evidence bar, no project enters production without naming the project it defunds. A greenlight that displaces nothing isn't a decision; it's an accumulation. The pet project has to name its victim in front of everyone. If the room can't say which kill pays for the yes, the project isn't greenlit. It's smuggled.
Scope note: licensed commitments, platform exclusives, and tax-credit teams can't defund each other; the rule governs the discretionary slate, where the smuggling lives.
Would the rule have saved my Friday kill? No; a channel that outranks the gate outranks the rule. It changes the price: the resurrection happens in daylight, named, in front of the room. Daylight is the one currency back channels can't spend.
AI is the multiplier. The greenlight is what you're multiplying.
Gameometry pairs operator judgment with AI to read slates faster and price their risks: gate sheets, evidence bars, kill reviews that hold when the room gets loud. If your slate is already thirty deep, that's the conversation. gameometry.io/contact