Every game you see on HKQS Games represents a small stack of decisions that, in aggregate, define what the platform is. Most of those decisions are invisible to players, which is a shame, because they are also the decisions that determine whether the catalog is worth your time. This piece is an attempt to make them visible — to walk through the process we use to evaluate, test, and approve games, and to be honest about the times that process has led us to say no.
I lead design at HKQS Games, which means I sit in most of the meetings where a candidate game gets discussed. The process below is not theoretical. It is the one we actually run, week after week, on every submission and every prototype we consider shipping. Where useful, I will point to broader industry conversations — like the ones collected at the Game Developers Conference archives — that have shaped how we think about curation.
The Starting Question
Before any evaluation rubric, before any test plan, we ask a single question of every candidate: would we be glad this game exists if we had not published it? It sounds soft, but it is the most useful filter we have. A game that is merely competent — that does nothing wrong but also does nothing you would miss — tends to fail this test. A game that has a real idea, even an imperfect one, tends to pass. The goal of the catalog is not to have everything. It is to have things that earn their place.
This question matters because the alternative is volume. A platform that ships everything that runs will accumulate a catalog of games that are fine, that load, that do not crash, and that no one will remember. That is the failure mode we are trying to avoid. Gamedeveloper.com's coverage of curation and discovery makes the same point from the business side: in an attention-scarce market, curation is the product, not an obstacle to it.
The catalog is not a bucket. It is an argument about what casual games can be. Every title we add is a claim; every title we reject is a claim we chose not to make. — Reza Pratama, design lead
The Curation Economics: Why Saying No Pays
The economics of curation are not obvious, because curation looks expensive in the short term and pays off in the long term in ways that are hard to put on a quarterly spreadsheet. The instinct, especially under growth pressure, is to ship more — more games mean more sessions, more sessions mean more impressions, more impressions mean more revenue. That instinct is correct as far as it goes, but it stops going far enough. According to Newzoo's mobile market reports, the global mobile games market reached roughly $92.6 billion in 2024, with the long tail of titles generating an increasingly small share of that revenue. The top 100 titles captured over 70 percent of mobile game spend, while the bottom half of catalogs collectively earned under 5 percent. Volume, in other words, no longer pays the way it once did.
This is the structural reason curation works. A catalog full of competent-but-forgettable games competes for the same small slice of long-tail attention that everyone else's catalog of competent-but-forgettable games is also competing for. A catalog of titles that have a real point of view — even if it is smaller, even if it loses some players who wanted breadth — earns a different kind of attention. Players who trust the catalog come back, recommend it to friends, and stay long enough that lifetime value compounds in a way that install-spikes from volume publishing never do.
Of evaluated candidate games made it to publication in the most recent quarter — roughly six out of ninety candidates. That is a lower rate than most platforms publish, and we think it is the point. Sensor Tower's publishing reports show that the average casual games platform ships between 30 and 60 percent of submitted titles. The gap between 7 and 50 percent is the gap between "curated" and "aggregated."
There is also a more subtle economic argument that is worth naming. Curation costs team time — review hours, test sessions, the back-and-forth with developers that turns an almost-good game into a good one. That time is real money. But the time is also what makes the catalog trustworthy, and trust is the asset that compounds. Every six months we have run the catalog, the cost-per-acquisition of a new player has trended down, not because we have gotten better at marketing, but because returning players bring new players with them. Think with Google's research on player retention covers the broader pattern: in casual gaming, organic word-of-mouth is the only acquisition channel whose cost goes down as volume goes up. Curation is the lever that pulls that channel into existence.
Volume publishing is a bet on a market where every game competes on price. Curation is a bet on a market where every catalog competes on trust. We are explicitly making the second bet, and we are willing to ship fewer games to win it. — Internal strategy memo, HKQS Games
The Four Gates
Once a candidate clears the starting question, it goes through four sequential gates. Each gate is a hard filter: fail any one and the game does not ship, regardless of how well it did on the others. This is deliberate. A game that is beautiful but broken, or fun but unfair, or fast but hostile to the player's time, is not a game we want to publish. The gates force us to be honest about which of those flaws we are willing to excuse — and the answer is usually none.
Gate 1: Does it teach itself?
The first gate is about onboarding, and it is the one that kills the most candidates. We put a first-time player in front of the game with no instructions, no tutorial prompts, and no help. If they cannot understand the core mechanic and complete a basic interaction within sixty seconds, the game fails. This is a stricter bar than most publishers apply, and we apply it because instant-play platforms live or die on the first minute. There is no install investment to keep a player around while they figure things out — if they do not get it, they leave.
Roughly 40 percent of candidate games fail this gate. Most fail not because their mechanics are too complex, but because they front-load complexity the player does not need yet. The fix is almost always subtraction — remove the menu, remove the tutorial, let the first screen be the first puzzle — but that fix has to come from the developer, and not every developer is willing to make it.
Gate 2: Is it fair?
The second gate is about fairness, and it is the one that most often produces arguments inside the team. A fair game is one where failure feels like the player's fault and success feels earned. An unfair game is one where outcomes feel arbitrary — where the player loses to a mechanic they could not have predicted, or where difficulty is created by withholding information rather than by raising the challenge. We test this by playing the harder tiers ourselves and asking, after each loss, whether we understood why we lost. If the answer is "I am not sure," the game is not ready.
This gate is where we filter out the most common monetization-driven design sin: difficulty spikes engineered to push players toward in-game purchases. A puzzle that becomes unsolvable without a paid hint is not a harder puzzle. It is a paywall wearing a puzzle's clothes. We reject these without hesitation, and we have turned down otherwise polished games because their late-game economies were structured this way.
Gate 3: Does it run on a cheap phone?
The third gate is performance, and it is the most measurable. Every candidate has to hit our three-second load budget and maintain a stable frame rate on a mid-range Android device. This is the gate our engineering team owns, and the one Sari writes about in detail in our piece on instant play performance. The short version: a game that runs beautifully on a 2024 flagship but stutters on a three-year-old phone is not a game we can ship, because most of our players are not on 2024 flagships.
This gate kills fewer candidates than the first two, but it kills them later in the process, which is more painful. A game that has cleared onboarding and fairness is a game we have often already fallen in love with. Telling a developer that their polished, fair, well-taught game does not run well enough on cheap hardware is one of the harder conversations we have. We have it anyway, because shipping a game that runs poorly for half our audience would be worse.
Gate 4: Does it respect the player?
The fourth and final gate is the values gate, and it is the one we cannot reduce to a metric. Does the game respect the player's time, attention, and dignity? Are the ads placed at natural breaks, or do they interrupt the action? Does the game use dark patterns — fake countdowns, guilt-tripping characters, manufactured urgency — to drive engagement? Is there anything in the experience that we would be embarrassed to have a friend see us playing?
This is subjective, and we accept that. It is also the gate that most reliably predicts whether a game will still feel good a month after launch. Games that pass the first three gates but fail this one tend to perform well for a week and then collapse in retention, because players do not come back to games that made them feel manipulated. The values gate is, in that sense, the long-term-business gate disguised as an ethics gate.
What the Rejections Look Like
It is worth being concrete about what does not make it through, because the rejections say more about the catalog than the approvals do. In the last quarter, we evaluated roughly ninety candidate games and shipped six. The eighty-four that did not ship broke down, broadly, into a few categories.
- Clones with no point of view. A perfectly competent water sort puzzle that is indistinguishable from four others we already have. We do not need a fifth. The bar for a clone is that it must do at least one thing meaningfully better than what we already publish, and most do not.
- Ad-economy games. Games whose design is shaped around ad placement rather than around play. You can usually tell these within five minutes: the natural break points exist, but they are padded out to fit ad slots, and the actual puzzles feel like filler between commercials.
- Polished but hollow. Games with beautiful art and smooth performance that have nothing to say. They pass gates one, three, and four, but they fail the starting question — we would not be glad they existed if we had not published them, because they do not do anything that has not been done better elsewhere.
- Almost-but-not-yet. The hardest category. Games with a genuine idea that is not yet realized — a clever mechanic that is not fully explored, a fair puzzle that needs another pass on difficulty, a beautiful game that needs another month of optimization. We try to send these back with specific feedback rather than a flat no, because the idea is worth saving even when the execution is not there yet.
Saying no to a game that almost made it is the hardest part of this job, and the most important. A catalog is built as much by the almost-good games you refuse to ship as by the great ones you do. — Design review notes, HKQS Games
Case Study: When a Game Failed at the Last Gate
One of the most useful examples we can share is a recent game that passed three of the four gates and failed at the fourth, because the failure mode is instructive. The candidate — call it Tower Stack Color for the purposes of this piece — was a stack-sort puzzle with genuinely beautiful art direction, smooth frame rates across our test devices, and a one-screen onboarding that taught the core mechanic in under thirty seconds. By Gate 1, Gate 2, and Gate 3 standards, it cleared comfortably. It failed at Gate 4, the values gate, and the story of why is worth telling.
The issue surfaced in our extended playtest, which is the protocol we run on every candidate that clears the first three gates. The protocol is simple: three team members play the game for ninety minutes each, across multiple sessions, and we look for anything that makes us uncomfortable. The Game Developers Conference talks on dark patterns in casual games are a useful reference here, because they catalogue the techniques the industry has converged on. Tower Stack Color used three of them.
- Manufactured urgency. A countdown timer appeared on each stack placement, presented as a difficulty mechanic. In practice, the timer existed to push the player toward the "freeze time" power-up, which was purchasable with premium currency.
- Guilt-trip NPC dialogue. A character appeared on level-loss screens with sad expressions and phrases like "I guess we are not a good team after all." The dialogue was specifically written to trigger a guilt-then-relief loop when the player bought a continue.
- Loss-aversion framing. The game displayed progress as "you have completed 80 percent, do not give up now" on exit screens — a framing explicitly designed to make leaving feel like failure, which is the textbook definition of a retention dark pattern.
Three team members, ninety minutes each, multiple sessions. The playtest is designed to surface patterns that the standard five-minute evaluation does not catch — specifically the dark patterns that emerge only after a player has invested enough time to feel the loss aversion. Gamedeveloper.com's coverage of dark patterns in mobile games documents the same point: most abusive design shows up only after the first session, which is why most quick-review processes miss it.
We sent the developer a list of the three patterns, the specific screens where they appeared, and a polite version of "we will publish this the moment these are gone." The developer came back with a counter-proposal: keep the urgency timer, but make it cosmetic rather than mechanical. We declined, because a cosmetic urgency timer still manufactures the same affective response in the player, just without a paywall at the end of it — which is arguably worse, since the player is being manipulated for engagement rather than for revenue. The conversation ended there. The game did not ship. We still think it could have been a good game. We also still think we made the right call.
The values gate is the one where you cannot hide behind a metric. It is also the one where, six months later, you are most likely to look back and be glad you made the call you made. Dark patterns work for a week and rot the catalog for a year. — Design review notes, HKQS Games
How the Six That Made It Got Made
The current catalog — Smash Blocks, Hunter: Evolve Uprising, Block Puzzle: Save Girl, Puzzle Hex, Puzzle: Water Sort, and Royal Match: King's Tale — each cleared the four gates in different ways, and the stories are instructive. Puzzle Hex almost did not ship because its hexagonal grid confused first-time testers; we worked with the developer to redesign the opening level so the mechanic taught itself, and it is now one of our best-retaining titles. Block Puzzle: Save Girl was rejected twice in earlier forms before the rescue metaphor was added, at which point the emotional anchor made the abstract shapes legible. Sky Thunder: Ace Strike took three optimization passes to hit our performance budget on low-end devices.
The pattern is that the games we ship are rarely the games we first saw. They are the result of a conversation — sometimes brief, sometimes months long — between our team and the developer, in which the four gates are not just pass/fail filters but a shared vocabulary for what needs to get better. The best developers welcome this. The worst treat the gates as obstacles to be argued around. We work with the first group and politely decline to work with the second.
Of evaluated candidate games made it to publication. That is a lower rate than most platforms, and we think that is the point. A catalog players can trust is built by being willing to say no to perfectly playable games that are not quite worth a player's time.
What Happens After Launch
Shipping is not the end of the process. Every game we publish is re-evaluated against the same gates on a rolling basis, using real player data. If first-session completion drops below our threshold, we look at whether an update broke the onboarding. If retention falls, we look at whether a difficulty change introduced unfairness. If performance regresses on low-end devices, the build gets rolled back. The gates are not just an entrance filter; they are an ongoing contract.
This means games occasionally come off the platform. It does not happen often, and it is never a surprise to the developer — the data has usually been telling the story for weeks. But it does happen, and we think the willingness to retire a title that no longer clears the bar is part of what keeps the catalog trustworthy. A platform that only ever adds and never removes eventually drowns in its own catalog.
How Industry Benchmarks Compare
It is useful to compare our curation numbers against the broader industry, because the comparison explains both why our catalog is small and why we think that smallness is the right answer. The most useful benchmarks we have are the public reporting from Sensor Tower and Newzoo, both of which publish quarterly snapshots of submission volume and approval rates across the major casual games platforms.
The numbers tell a story of widening divergence. The major mobile app stores approve roughly 40 to 60 percent of submitted games, with the long tail composed largely of reskinned clones and ad-economy titles that never meaningfully break out. The major instant-play platforms — the segment we are in — publish at higher rates, between 60 and 80 percent, because the bar to clear is mostly technical rather than editorial. Our 7 percent publication rate is roughly an order of magnitude lower than the comparable instant-play benchmarks. That is not a number we are embarrassed by; it is the number we are deliberately trying to produce.
Our publication rate versus the typical instant-play platform approval rate, per Sensor Tower and Newzoo tracking. The gap is the visible signal of the editorial process — a 7 percent rate means we are saying no to roughly 13 games for every 1 we ship. The cost is a smaller catalog. The benefit is a catalog that is trustworthy enough to recommend without caveats.
The other benchmark worth comparing is first-session retention, which is the metric we use most often to predict whether a shipped game will still be earning its place in three months. According to Newzoo's mobile retention benchmarks for 2024, the median free-to-play casual game loses roughly 65 percent of players within the first session, and 80 percent by day seven. The median title on our catalog loses roughly 38 percent within the first session and 52 percent by day seven. We do not attribute that gap entirely to curation — game quality matters most, and our developers deserve the credit there — but the curation floor does filter out the titles most likely to lose players in the first session, which is the segment most responsible for the industry's weak retention numbers.
The honest version of this comparison is that we are not chasing a different market. We are chasing the same market — casual puzzle players, the broadest audience in gaming — and we are simply refusing to ship the long tail of titles that has historically made casual gaming feel disposable. The bet is that the audience for a curated catalog is large enough to sustain a platform, and that the audience for a non-curated catalog is large but commoditized. Think with Google's player-trust research supports this view: in their 2024 survey of casual gamers, 71 percent of respondents said they had uninstalled a game within the first session because it felt manipulative or low-quality, and 64 percent said they had stopped using a games platform entirely because they no longer trusted the catalog. Trust, in casual gaming, is not an abstraction — it is a retention metric.
The industry's retention problem is, fundamentally, a curation problem dressed up as a UX problem. You cannot design your way out of a catalog full of games that were never worth a player's time in the first place. — Reza Pratama, design lead
The Future of Game Curation
It is worth being explicit about where we think curation is going, because the next few years are going to put pressure on every choice we have described so far. Three trends are converging on the casual gaming category, and each of them changes the math of curation in a way we have to think through ahead of time rather than react to later.
The first is AI-generated content. The cost of producing a competent-looking casual game has fallen dramatically, and it is going to fall further. Gamedeveloper.com's coverage of generative tools in casual game production estimates that the time-to-prototype for a stack puzzle or color-sort variant has dropped from weeks to days, with much of the asset production now automatable. The implication for curation is direct: submission volume is going to rise, and a larger share of those submissions are going to be "competent but hollow" — games that look fine and play fine and have nothing to say. The four-gate process is going to have to get faster, and the starting question is going to have to get sharper, because the volume of "fine" games is going to overwhelm any process that depends on slow editorial review.
The second is the gradual obsolescence of the device advertising identifier, which is the technical backbone of the current ad-economy model. Think with Google's work on Privacy Sandbox tracks this transition in detail: Apple's App Tracking Transparency has already cut identifier availability on iOS to under 25 percent of users, and Google's ongoing Privacy Sandbox work is moving Android in the same direction. As behavioral targeting becomes technically harder, the economic case for the ad-economy game design — the dark-pattern-driven, retention-by-manipulation model — weakens. That is good news for the kind of games we want to publish, and bad news for the kind we reject. Curation, in that future, becomes more rather than less valuable, because the games that survive without behavioral ad revenue are the games that earn attention by being good.
Our internal forecast for candidate submission growth over the next 24 months, driven primarily by generative tools lowering the cost of casual game production. The Newzoo 2024 market report projects similar growth in casual game releases globally. The implication for curation is that the cost of saying no — measured in team hours per evaluated candidate — has to come down, even as the bar for saying yes has to stay the same.
The third is platform-level curation itself becoming a feature players search for. The era of "everything is available somewhere" has produced, as a predictable side effect, an exhaustion with choice. Sensor Tower's user-survey work has been picking up this signal for the last two years: a growing share of casual players report that they discover new games through specific platforms or curators they trust, rather than through app-store search. That is the trend we are betting on. A platform that has earned the right to be a discovery surface — because it has demonstrated a willingness to say no — is worth more in that world than a platform that has merely accumulated volume.
None of these trends are reasons to relax the gates. All three are reasons to keep them sharp. The future of curation is not a softer version of the present; it is a more deliberate one, in which the willingness to reject becomes the most valuable thing a platform can offer. We will keep running the gates the way we run them now, and we will keep being honest about the times we get it wrong, because the only way to earn the right to curate is to be willing to defend the calls — including the ones that, in hindsight, we wish we had made differently.
Curation is a long bet. The platforms that win the next decade of casual gaming are not going to be the ones that published the most games. They are going to be the ones that published the fewest games worth trusting. — Reza Pratama, design lead
Why We Do It This Way
The easy version of running a games platform is to publish everything that runs and let the market sort it out. That model scales, and it is what most large platforms default to. We do not use it, because the cost of that model is paid by the player — in time lost to bad games, in attention drained by manipulation, in the slow erosion of trust that makes every future game harder to discover.
The model we use is slower, more expensive, and harder to defend on a spreadsheet. It requires a team that can play games critically, give specific feedback, and have hard conversations with developers. It requires saying no far more often than we say yes. And it requires believing, against a lot of market pressure, that a small catalog of games you would genuinely recommend is worth more than a vast catalog of games you would not. That is the bet we are making every time we run the gates. So far, the players seem to agree.