Climb Test Prep app icon

Climb

Adaptive Test Prep · iOS

An AI tutor who understands your student. Built around mastery, not engagement.

Adapts to your student's mistakes Native to the digital SAT format Wren reads handwritten work A season for about one tutoring hour Data stays on your device

A KyrosWorks product

60-second tour · Click for sound

How it works

Set a target. Climb builds the plan.

You tell Climb your target score and test date. It runs a diagnostic, commits to a plan, and adjusts the cadence to however much runway you have — real progress on 15–20 minutes a day for the typical student, more if the test is close, less if you started early. The application is structured as a sherpa-guided climb to your target on the SAT® exam — the metaphor mirrors a real Mt. Everest base-camp trek the founder is taking with his son in 2027.

Climb onboarding screen asking for target SAT score and test date
1

Size up the climb

What should Wren call you, your target score, your test date, your most recent score if you have one. That's it — none of it is locked in, and you can change it later.

Nothing else to fill out. Climb starts working from four numbers, not a twenty-field survey.

Climb plan screen showing target score, test date, days remaining, and the three-stage study approach
2

The diagnostic, then a real plan

The first 25–50 questions are quiet on purpose — Climb is watching what trips you up, where you're solid, where you rush, before it commits to anything. Once the pattern is clear, it names the highest-value subtopic, walks you through a short lesson, drills it, and checks that it stuck before moving on.

"First 25–50 questions, I'm mostly quiet — watching what trips you up, where you're solid, where you rush." — Wren, at plan start

Climb practice question screen with multiple-choice answers, correct-answer confirmation, and a worked explanation
3

Practice that targets weak topics, not random ones

A Beta-posterior mastery model tracks what your student actually knows per subtopic. Leitner spaced repetition resurfaces cards right before they'd be forgotten. Bootstrap propagation transfers learning across difficulty tiers, so a correct answer at the Hard level also raises confidence at Medium.

A missed Easy comma-splice yesterday returns today. Nail it at the Hard difficulty and confidence at Medium rises too.

Climb today's plan screen showing a mix of review, new, and sharpen questions sized to the day
4

A daily cadence sized to your test date

Every session opens on today's plan — a mix of review, new material, and sharpening, sized to the days left before your SAT. When the queue is empty, "0 due today" shows as a feature, not a failure. No streak, no penalty for a missed day.

Typical cadence: 15–20 minutes a day, roughly five days a week. Climb compresses that automatically as test day gets close.

Your AI tutor

Wren teaches against your actual mistakes

Wren isn't a chat window bolted onto a quiz bank. She reads your student's recent attempts and per-topic mastery before saying anything, anchors explanations in concrete past moments, and remembers what tripped him up weeks ago — the compounding effect that makes practice actually add up instead of resetting every session.

Wren chat screen introducing herself, explaining the diagnostic and how she reads photos of handwritten work

Memory across sessions, three teaching modes

Explain (default) — names the trap and the move in three to five sentences. Diagnose — "walk me through how you got there," only when a wrong answer suggests a deeper misconception, so the student surfaces his own reasoning before the explanation lands. Reassure & Reset — after three wrong answers in a row, Wren stops teaching, acknowledges the struggle, and offers a topic he's strong in instead of piling on.

"You got the parallel-structure version yesterday. Same move here. Pick the option whose verb form matches the first item, not the one that sounds smoothest."

[SCREENSHOT NEEDED]
Wren reading a photo of
handwritten work and pointing
to the exact wrong step

She can read a photo of handwritten work

Most of the SAT still happens on scratch paper. Snap a photo of the work in the Wren tab and she reads it — not just the final answer, the actual steps — and points to the specific line where the error happened. That's the difference between "you got it wrong" and "this is where the sign flipped."

The photo never leaves the device beyond the request itself — Climb doesn't store or serve it back from a server. See privacy.

The climb

Real progression, real mountain

Mastery growth maps to altitude on Mt. Everest — Kathmandu at the start, Everest Base Camp once fundamentals are secure, Camps 1–4 and the summit at advanced mastery. Every answer writes a real altitude time-series. It's the same route the founder is walking with his son in 2027 as a graduation trip. No fanfare at stage transitions, no achievement sounds — the mountain is real, so the metaphor stays earned.

[SCREENSHOT NEEDED]
Course / mastery view —
altitude progress up the
Everest route

Altitude is earned, not unlocked

There's no level-up animation for opening the app. Altitude moves only when mastery moves — a subtopic checked off, a Leitner card retired for good. The view is honest about where a student actually stands relative to test day, not where a streak counter says they should be.

Why we don't gamify

The easiest features we never built

Climb's first user was the founder's son, a kid whose attention was trained by Fortnite, Apex, and Genshin Impact. Every product instinct said to meet him there: XP for right answers, a streak counter, a level-up chime. That layer would have taken a weekend to build, and every engagement metric would have looked better for it.

We read the research before writing the code. A 2023 meta-analysis of educational gamification found the standard points-badges-leaderboards trio does not measurably improve a student's competence. Worse, the over-justification effect (Lepper, 1973) shows that bolting extrinsic rewards onto studying can crowd out whatever real interest was forming. Streaks reward opening the app. The SAT does not score attendance.

  • No XP
  • No streaks
  • No daily quests
  • No badges
  • No levels
  • No character classes
  • No leaderboards
  • No engagement notifications
  • No subscription trap
  • No engagement-loop manipulation

We considered each of these mechanics, audited it against the research, and removed it on purpose. Each removal is documented in the research below. The rule that survived: every reward loop must be denominated in real mastery gains, not presence. If a mechanic rewards opening the application, it is wrong. If it rewards getting better, it stays. What your student feels instead of a dopamine loop is a plan that visibly shortens the distance to their target score. That turns out to be its own motivation, and it is the one that survives contact with test day.

Content

A full question bank, native to the digital SAT

Every question runs in the same digital-first format the actual exam uses — digital-SAT practice, not paper-test drills. Coverage spans all four Math domains and all four Reading and Writing domains the digital SAT tests, plus adaptive lessons Wren generates for the exact subtopic a student needs, not a static curriculum.

Climb intro screen: your plan not generic prep, practice that compounds, Wren your AI tutor

Math

Algebra
Advanced Math
Problem-Solving & Data Analysis
Geometry & Trigonometry

Reading & Writing

Craft & Structure
Information & Ideas
Standard English Conventions
Expression of Ideas

Question count: [QUESTION COUNT TBD] — test-eligible questions across all eight official digital-SAT domains, each with a baked-in worked explanation. Every wrong answer opens a review path back through Wren — a short lesson on the specific rule, drill questions, and a check before the plan moves on.

Privacy

Your student's work stays on their device

Climb is built with on-device data sovereignty as a default, not an add-on.

📱

Study data lives on-device

Answers, mastery records, and progress are stored locally. None of it reaches our servers.

💬

Wren's chat history is local

Conversations with the AI tutor are persisted on the device, not in a cloud account tied to your student.

🖼️

Handwriting photos stay on the phone

Photos of handwritten work are used to generate Wren's reply and are not stored on our servers.

🚫

No data sale, no telemetry

We don't sell data and don't run analytics or tracking SDKs. No third parties beyond Apple (billing) and Anthropic (the AI itself).

On the server side we hold the bare minimum to run the subscription: an App Attest device identifier (Apple's proof that a request comes from a real Climb install), current subscription state, and a per-device monthly usage counter. Full privacy policy.

Pricing

Priced for a test season, not a forever app

One AI-tutoring tier, billed by Apple. Cancel auto-renewing plans anytime in iOS Settings and keep access until the period ends.

Monthly

$19.99/mo

Auto-renewing

Cancel anytime. Best if you want flexibility over commitment.

Annual

$149/yr

Auto-renewing · ~$12.42/mo equiv.

Best for multi-cycle prep — PSAT, SAT, and a retake in the same subscription.

The subscription gates Wren, the AI tutor, up to a generous monthly usage cap — typical study usage runs well under it. The adaptive practice engine, the full question bank, and lesson content work the same on every tier.

Why Climb

How it stacks up

Acely, R.test, and Bonsai roughly share Climb's price range and all layer some kind of AI on top of SAT content. Khan Academy and Khanmigo are free and excellent for content depth. Here's where Climb is actually different — not a claim that it beats every tool on every dimension.

What matters Climb Generic AI SAT apps Khan / Khanmigo
Remembers mistakes across sessions Yes — anchors new explanations in past ones Varies; many reset context each session Not designed around cross-session memory
Reads a photo of handwritten work Yes — points to the specific wrong step Uncommon Not a core feature
Native digital-SAT format Yes, every question Varies by app Question bank not digital-SAT-native
Data stored on-device by default Yes Typically cloud-first Cloud account required
Gamification (streaks, XP, badges) Deliberately none — see research Common Points/badges present in places
Price $59 for a 4-month season Roughly $20–50/mo, ongoing Free
Content depth & polish Growing digital-SAT-native bank Varies Very deep, years of investment

Khan Academy and Khanmigo have a content library Climb doesn't try to match on depth — Climb doesn't replace them there. What Climb adds is a tutor built around one student's actual error pattern, priced and structured for a single test season rather than an open-ended subscription. Your student can paste a question into any general chatbot and get an answer; the gap is a tutor that already knows what tripped him up three weeks ago and picks the right teaching move for this exact moment.

Measured

Even our worst run beats every raw frontier model

We benchmarked Climb-Wren against five raw frontier AIs: Anthropic's claude-sonnet-4-5 (Climb's model), Anthropic's claude-opus-4-7 (the flagship — most expensive model on the market), OpenAI's gpt-5 (behind ChatGPT), and Google's gemini-2.5. Four independent eval runs. The numbers below report each condition's worst single run — an anti-cherry-picked floor. If a parent or competitor re-runs the benchmark, they will mathematically land at or above what we publish here. Read the full methodology →

Shape-appropriate teaching · regex-scored · 17 mode scenarios · worst of 4 runs

2.65 vs 0.35

Out of 3. Climb-Wren's worst run versus raw Opus 4.7's worst run. Opus is Anthropic's flagship and the most expensive frontier model on the market. Climb beats it by a factor of 7.5×, on a model class that costs 5× less. The same comparison versus raw GPT-5 (worst 0.53) and raw Gemini-2.5 (worst 0.53) is +2.12. Raw Opus alone is the worst of the four raw models on this dimension — the most expensive model is not the best at unaided tutoring shape.

Mode-shape · un-fakeable

2.65 / 3 worst of 4 runs

+2.12 vs raw-Sonnet (0.53) · +2.30 vs raw-Opus (0.35) · +2.12 vs raw-GPT-5 (0.53) · +2.12 vs raw-Gemini (0.53)

Pure regex over the reply text. Worked example? Numbered steps. Diagnose? Ends with a question. Acknowledge? ≤ 2 sentences, no follow-up. Reassure-reset? Opens with a pause cue. No LLM judge — no judge bias. The most defensible single number we have, and the one we lead with.

Pedagogical quality

1.86 / 3 worst of 4 runs

+0.54 vs raw-Sonnet (1.32) · +0.91 vs raw-Opus (0.95) · +0.77 vs raw-GPT-5 (1.09) · +0.86 vs raw-Gemini (1.00)

Is the teaching move appropriate to the moment? Whether the student walks away understanding something they didn't a minute ago. LLM-judged against a rubric — disclosed because judge-mediated scoring has bias risk. The 4-run floor framing absorbs judge variance.

Error localization

1.41 / 3 worst of 4 runs

+0.18 vs raw-Sonnet (1.23) · +0.27 vs raw-Opus (1.14) · +0.37 vs raw-GPT-5 (1.04) · +0.41 vs raw-Gemini (1.00)

Does the reply name the root-cause subtopic? Climb wins here despite deliberately staying quiet on acknowledge and reassure-reset turns — modes where naming the subtopic is the wrong move.

Climb-Wren on Opus — the ceiling

3.00 / 3 every run perfect

Mode-shape 3.00 / 3 across all 4 runs · Pedagogy 2.09 (worst) → 2.23 (best)

When we route Wren's scaffolding through Opus 4.7 instead of Sonnet 4.5, mode-shape is perfect on every single run. Pedagogy lifts another ~0.1 above Wren-on-Sonnet. The system is what wins — but it scales with the brain. We ship on Sonnet because the gap is small and the price is 5× lower. A premium tier remains an option.

How the measurement works

Twenty-two scenarios — seventeen of them tagged with a teaching mode (worked example, diagnose, explain, acknowledge, reassure-reset), five generic. Three student-history fixtures (forgotten-mastery, baseline-mid, distractor-trap-rusher). For each scenario, six assistants reply: Climb-Wren on Sonnet, Climb-Wren on Opus, plus raw-Sonnet, raw-Opus, raw-GPT-5, and raw-Gemini (no system prompt, no tools, no engine recommendation — just the user message). The Climb-Wren paths also get full session history, the engine's mode recommendation, and tools wired to the database.

Four scoring dimensions per reply. One is pure regex (mode-shape) and immune to judge bias. Three rely on an LLM judge against a written rubric — we disclose this. To absorb judge variance and provider non-determinism, the full eval is run four independent times and we publish the worst per-condition score across runs (4-run floor). A skeptic re-running the benchmark lands at or above these numbers.

The system around Climb-Wren is what changes — not the model underneath. Climb ships on Sonnet 4.5. Across all four raw brand baselines and all four runs, the system wins.

Reproducible: swift run ClimbEval --live against the open-source ClimbEval harness. Live API cost per full 6-condition × 22-scenario run: ~$3. Total spend for the 4-run aggregate: ~$13.

What we don't claim

  • n = 17 mode scenarios × 4 runs is a real sample size, but not a meta-analysis. The worst-of-4 framing is anti-cherry-picked: a skeptic running the eval lands at or above what we publish.
  • The control conditions are "frontier model alone," not "ChatGPT.com's actual UI." Your kid can paste a question into ChatGPT and get a response. They just won't get one that knows what they got wrong three weeks ago and chose the right teaching shape for this exact moment.
  • Pedagogical-quality and error-localization scores rely on an LLM judge. Mode-shape (the headline) is regex-scored with no judge involvement — that's where Climb's margin is widest and most defensible.
  • Gemini was tested on the 2.5-Flash tier (Pro requires Tier 1 billing promotion that's still propagating). Flash already loses to Climb by +2.12; Pro is the larger model and unlikely to close the gap on shape-appropriate teaching, but we'll re-publish with Pro once it's accessible.
  • Climb-Wren-on-Opus is the absolute ceiling (perfect mode-shape every run), but we ship Climb on Sonnet because the marginal gain at 5× the cost doesn't justify a higher product price for the average student. We'd rather your kid have Wren than a 5× markup for a ceiling effect.
  • Climb-Wren scored fractionally below raw-GPT-5 on prior-context recall in some runs. Disclosed; not hidden. Climb's edge isn't in name-dropping past episodes — it's in picking the right teaching shape and finding the right anchor at the right moment.

Why these design choices

The research base

Climb's design is grounded in the educational-psychology literature on what actually moves learning outcomes. Every claim below links to the underlying work.

One-on-one tutoring is the gold standard. Climb's Wren approximates it.

Bloom (1984) found that one-on-one tutoring produces a two-standard-deviation improvement over conventional classroom instruction. The mechanism is not more time — it is that the tutor models the student's mental state, identifies the specific misconception, and intervenes against that misconception directly. Wren does this through persistent memory, mastery-state queries, and diagnostic questions that ask the student to surface his reasoning before the explanation lands.

Bloom, B. (1984). The 2 Sigma Problem. Overview · Carnegie Learning's Cognitive Tutor lineage: Anderson, Corbett, Koedinger & Pelletier — Cognitive Tutors: Lessons Learned.

Mild confusion correlates with learning gain. The tutor's job is not to dissolve it.

Kapur's productive-failure work shows that students who struggle with a problem before receiving instruction outperform students given direct instruction first — by effect sizes equivalent to two to three years of typical schooling on transfer tasks. Wren's diagnostic mode is engineered around this finding. When she asks "walk me through how you got there," she sustains productive disequilibrium long enough for the student to construct the right concept himself, rather than resolving the confusion immediately.

Kapur, M. — Productive Failure (overview). Roediger & Karpicke on the testing effect: The Power of Testing Memory (2006).

Spaced retrieval beats massed cramming, by substantial margins.

Bjork's "desirable difficulties" framework, supported by decades of cognitive-science replication, shows that spaced retrieval practice produces durable learning while massed cramming produces short-term recognition that does not survive the actual test. Climb's Leitner queue is a working spaced-repetition implementation; the adaptive sampler weights against recency to enforce spacing.

Bjork & Bjork — Making Things Hard on Yourself, But in a Good Way. Background: Desirable difficulty.

Extrinsic rewards can crowd out intrinsic motivation. So we do not add them.

Lepper, Greene & Nisbett (1973) demonstrated the over-justification effect: bolting extrinsic rewards (XP, badges, points) onto an activity in which a student might develop intrinsic interest reduces that intrinsic motivation when the rewards are removed. The implication for an app preparing students for the SAT® exam is direct: if we want a student to develop genuine interest in mastering the material, layering XP and badges on top is the most counterproductive thing we can do. Self-determination theory (Deci & Ryan) names the actual drivers of sustained motivation — autonomy, competence, relatedness — none of which are points.

Lepper, Greene & Nisbett (1973): Undermining children's intrinsic interest with extrinsic reward · Ryan & Deci (2000): Self-Determination Theory and the Facilitation of Intrinsic Motivation.

The standard gamification trio does not measurably improve learning competence.

A 2023 meta-analysis of educational gamification found that points, badges, and leaderboards — the most-borrowed mechanics in ed-tech — do not measurably improve felt competence. What does move competence: appropriate-difficulty challenges with visible feedback. That is precisely what Beta-posterior mastery, Leitner spacing, and Wren's anchored explanations deliver, without badges.

2023 meta-analysis of educational gamification — Educational Technology Research and Development.

Push notifications reduce learning performance. So Climb does not send any.

A 2022 study found that mobile push notifications, even non-engagement-oriented ones, measurably reduce learning performance. Climb sends no push notifications. The spaced-repetition queue creates a natural daily rhythm based on memory science; we do not need to interrupt the student to drive return visits.

Push notifications and learning performance — Computers & Education: Open, 2022.

For parents

What you'll want to know

Read carefully. Fact-check anything. The answers below are honest, including the limitations.

How is this different from Khan Academy, Acely, R.test, Bonsai, Magoosh, or UWorld?

Khan, Magoosh, and UWorld have excellent content libraries — Climb does not try to replace them on content depth. Acely, R.test, and Bonsai are closer peers on price and AI-tutoring positioning. What Climb adds: a tutor who remembers your student's specific struggles across sessions, reads photos of handwritten scratch work and points to the exact wrong step, and stores study data on-device by default rather than in a cloud account. Full comparison above. Khan's points-and-badges layer is one of the elements Climb deliberately does not replicate, for the reasons in the research above.

How much screen time will my student spend in this?

Exactly as much as is mastery-productive, and no more. The application does not engineer engagement. There is no streak shaming him into opening it daily, and no notification pulling him back in the evening. The Leitner queue surfaces cards on a memory-science cadence; when the queue is empty, "0 cards due today" is shown as a feature, not a failure. Typical use is fifteen to twenty minutes per session, most days of the week.

Is my student's data safe?

Yes. Study progress, answers, and chat history with the tutor are stored only on the device — none of that ever reaches our servers. AI requests, including handwriting photos, flow from the device through our Cloudflare Worker proxy to Anthropic; the proxy is a passthrough and does not log or store the content of the requests. On the server side we hold the bare minimum to run the subscription: an App Attest device identifier, current subscription state, and a per-device monthly usage counter. No telemetry. No analytics. No third parties beyond Apple (billing) and Anthropic (the AI). Full privacy policy.

What does the AI tutor actually do, in plain terms?

Three modes. Explain — when your student gets a question wrong on something he mostly knows, Wren names the trap and the move in three to five sentences. Diagnose — when the wrong answer suggests a deeper misconception, Wren offers: "If you want, walk me through how you got there — otherwise I can show you the move directly." When he opts in, his answer reveals the specific misconception, and Wren teaches against that directly. Reassure & Reset — when he has missed three questions in a row, Wren stops teaching, acknowledges the struggle, and offers to switch to a topic he is strong in, to rebuild confidence. She can also read a photo of handwritten scratch work and point to the exact step where the error happened. Mode selection is gated on numerical mastery thresholds and recent-attempt patterns, not on intuition.

What does it cost?

Three tiers, all billed by Apple: Monthly at $19.99/month, cancel anytime. Season Pass at $59 for four months — a one-time, non-renewing charge sized to cover signup through a typical test date, and the option we'd recommend for most students. Annual at $149/year for students prepping across multiple cycles (PSAT, SAT, a retake). The adaptive practice engine, the full question bank, and lesson content are the same on every tier — the subscription gates Wren, the AI tutor, up to a generous monthly usage cap that typical study usage runs well under.

Who is this not for?

Climb works best for students with at least an 1100 floor who can mostly self-direct. If your student needs structured tutoring through fundamentals (reading comprehension, basic algebra), this is not the right tool yet — the application assumes the substrate is in place and works on closing the gap to a target score. To be direct: Climb is not a replacement for a human tutor when one is needed, and it is not a substitute for the accountability scaffolding some students need to study at all.

How early-stage is this?

Early. Climb is currently in TestFlight (iOS beta) while we finish App Store review. The founder built it for one user — his own son — and is opening it up as that work proves out. If you would like access, get in touch and we will discuss whether it is a fit for your student before adding you to the test group.

The story

Built for one student first

The founder's son, Westley, scored 1140 on his first SAT® exam. Target: 1340. Eighteen weeks of runway.

Westley plays Fortnite, Apex, Genshin Impact, and Baldur's Gate 3 — his attention pattern is game-trained. He expects rich feedback loops and felt progress. The temptation was to graft those mechanics onto the application: XP for every correct answer, streaks for daily use, character-class progression, all of it.

The honest move was different. The research is clear that those mechanics improve return-visit metrics but do not measurably improve learning. Westley deserves a tool that actually teaches him — not one that engineers his attention to look like it is teaching him.

So Climb went the other direction. Real adaptive practice. A real AI tutor with memory. A real mountain — because the founder and Westley are walking the Everest Base Camp trek in 2027 as Westley's high-school graduation gift, and the application's progression mirrors the literal route they will walk together. The climb is real.

If it works for Westley, it can work for other students. That is the path.

About

From KyrosWorks

KyrosWorks is a one-person operation building AI-native software for the kind of problems people reliably forget to solve, sustain, or get right by hand. Climb is one of those products. Open Road Ace — a millisecond-accurate pacing application for open-road time-trial racing — is another.

Each KyrosWorks product begins with a problem the founder personally has, is validated with users who share that problem, and ships when it works under the conditions where it actually fails. That is the methodology.

Get TestFlight access

Climb is on TestFlight while App Store review finishes. Send a note about your student (current SAT® score, target, timeline, anything else worth knowing) and we will discuss whether it is a fit before adding you to the test group.

will.hays@gmail.com

In the interest of transparency: Westley was Climb's first confirmed test user. The project is early. We would rather work with a small group of motivated students than a large group of curious ones.