A Cabo player who joined 3 days ago had a PQ of 1600. A veteran who’d been playing for 6 months had 1674. The entire point of PQ was to distinguish skill levels. When a 3-day player is within reach of a 6-month veteran, PQ is not doing its job.
The problem: our scoring formula was inflated. Every pillar — skill, consistency, volume, diversity — was too easy to max out. The result was a compressed cluster where everyone was “good” and nobody was “great.”
The original formula
PQ was computed as a composite of four pillars:
- Skill: Wilson lower-bound win rate (for Cabo) or average XP per session (for solo games), normalized to
[0, 1]. - Consistency: days active last 30 / 30 + longest streak / 30.
- Diversity: Shannon entropy of game distribution + breadth (games with ≥10 plays / total games).
- Volume: log10 of weighted XP.
The composite was: PQ = 1000 + 1000 × (weighted sum). Each pillar could reach 1.0, and the weighted sum could reach ~0.8 for a moderate player, giving PQ ~1800. Too high.
What went wrong: the five constants
1. Skill references were too low
Solo game skill refs (the denominator for average XP per session) were set at 120 XP. But the average Trek session earns about 50 XP. With a ref of 120, that’s 50/120 = 0.42 — already near the ceiling after just a few sessions. We raised solo refs 3× (Trek to 270, Recall to 360, Wordgrid to 500), making skill harder to max.
2. Consistency maxed out in 30 days
The original formula: daysActive/30 + streak/30. Playing every day for a month gave you 1.0 + 1.0 = 2.0, clamped to 1.0. Even 15 active days hit 0.5. We changed it to 0.4 × (days/60) + 0.6 × (streak/365). Now it takes 60 active days and a year-long streak to max out. A 30-day player with a 30-day streak earns ~0.35 (was ~1.0).
3. Volume grew too fast
log10(weightedXP/1000) gave 0.75 at just 5,600 XP — about two weeks of active play. We replaced it with a power function: (weightedXP / 10M)^0.3. Now 10K XP gives 0.06 (was ~0.6), 100K gives 0.16, 1M gives 0.50. Volume is the slowest-growing pillar.
4. Diversity threshold was too low
Playing a game 10 times counted toward diversity breadth. We raised the threshold to 50 plays. Playing Trek 10 times shouldn’t meaningfully count as “diverse.”
5. The composite was too generous
The original PQ = 1000 + 1000 × composite could give a composite of 0.6 for a moderate player, yielding PQ 1600. We changed it to PQ = 1000 + 1000 × composite². Squaring the composite means doubling effort only adds ~41% more PQ earned. Growth is deliberately slow.
The result
| Player profile | Old PQ | New PQ |
|---|---|---|
| 3-day casual | 1594 | 1026 |
| 1-month regular | 1633 | 1042 |
| 6-month daily | 1674 | 1121 |
| 1-year hardcore | 1721 | 1172 |
| 2-year completionist | 1810 | 1255 |
The spread went from 1594–1810 (216 points) to 1026–1255 (229 points). But the domainis now anchored at 1000 with deliberate slow growth. A 3-day player at PQ 1026 is clearly “just started.” A 2-year player at 1255 has earned their higher PQ through genuine long-term engagement and skill.
Why squaring the composite works
The composite score ranges from 0 to 1. A moderate player might have a composite of 0.4. Without squaring: PQ = 1000 + 400 = 1400. With squaring: PQ = 1000 + 160 = 1160.
The square makes PQ growth sub-linear. The first 30% of effort gives you PQ 1090. The last 30% (from 70% to 100%) gives you an additional 91 points. Diminishing returns. The best players are differentiated, but the ceiling is far away.
The daily rollup optimization
Originally, every session write triggered a Firestore function that recomputed the user’s PQ. At 825 sessions per day (275 DAU × 3), this meant ~8,250 Firestore reads per day. We replaced it with a daily cron job at 1 AM IST that recomputes PQ for all users who played that day.
Effect: ~1,650 reads per day — a 5× reduction. PQ values are updated once per day, which is fine for a score that changes by single-digit points per session.
What I’d Tell My Past Self
- Test with real data, not hypotheticals. Our original constants were calibrated against hypothetical players. Real players earned XP faster than expected.
- Squaring the composite is the simplest inflation fix. It’s one line of code and it’s mathematically grounded — diminishing returns.
- Raise references aggressively. Solo game skill refs at 120 meant average play scored 0.42. At 360, average play scores 0.14. That’s the right range.
- Volume should be the slowest-growing pillar. It’s the easiest to grind. Make it grow like
x^0.3, notlog10(x). - Batch recomputation is 5× cheaper than per-session triggers. PQ barely moves per session. Daily recomputation is sufficient.
