Ashish Noel · Published 1 August 2026
How Prelimo's questions are made
Prelimo's daily questions are generated by AI. This page says exactly how, what each question has to survive before it reaches you, and what happens when one is still wrong.
What the machine reads before it writes anything
Nothing is generated from a blank page. Two bodies of material sit under every question.
The first is the papers. Eleven real UPSC Prelims papers, 2015–2025, a little over a thousand questions, almost every one of them tagged against a fixed vocabulary of 207 concepts across the 12 General Studies subjects, and mapped further down to 800 subtopics drawn from the contents pages of the standard books. That tagging is what lets the system ask a useful question before it writes one. When UPSC has come at this concept before, what shape did it take, and how often.
Some of that corpus is now open to read. We are working through it question by question, writing our own explanation for each one and putting it through the same key-withheld check the daily questions face, and publishing the ones that clear it: past-year questions with answers and explanations.
The second is the news. Every night a job pulls the day’s current affairs and classifies each item against those same concepts. Most of the news maps to nothing on the syllabus and gets dropped, and that is the point of the step. A concept the day’s news happens to sit on top of is worth a question. A big headline that touches nothing in the syllabus is still noise.
What happens overnight
In the small hours, IST, the pipeline picks the day’s angles. Not the biggest stories. The ones sitting on top of a static syllabus concept the paper has a history of asking about. This is the static-dynamic nexus, and it is the whole product thesis in one line. Current affairs are not a separate subject to be memorised. They are the surface UPSC uses to ask a static question.
For each of the ten slots the generator drafts more than one candidate, in the statement-heavy formats the paper has moved to since 2023. How many of the above statements are correct. Assertion and reason. Match the pairs. Not the two-line one-liners a lot of question banks are still full of.
What every question has to survive
A question you see has cleared all of this. One that fails any of it is not softened and re-used, it is dropped.
Check 01
Shape, then repetition
A draft has to parse into a real question first: a stem, four options, exactly one key, an explanation for every option. Then its meaning is embedded as a vector and compared against the whole existing bank and against the last seven days of daily sets, so the same question cannot come back inside a week wearing different words.
Check 02
An independent correctness judge
Every candidate is re-solved from scratch by a second model with the answer key hidden from it. If that solver confidently lands on a different option than the key claims, the candidate is thrown out. Only a confident disagreement kills it. A checker that also failed on its own flakiness would quietly empty the set, which is its own kind of harm.
Check 03
Best of several, not first past the post
Candidates that survive the judge are compared head to head against a rubric. How good the trap is, how tightly the question is anchored to the news item it came from, whether each statement stands on its own rather than leaning on the one above it. The winner ships. The loser goes to a reserve bank rather than the bin.
Check 04
A search-grounded fact check
This is the layer for a question that is well formed, judged correct, and simply not true. Each factual claim is split into individual items, and a checker with live web search verifies them one at a time, with the answer key withheld from it. Then plain code, not a model, compares what the web said against what the question needs each item to be.
A statement meant to be false that is false is a working trap, so it passes. A statement meant to be true that turns out false breaks the key, so the question dies. Anything the web cannot verify never kills a question on its own.
Check 05
A second opinion from a different model family
A generator and a judge from the same model family can share a blind spot, and on India-specific constitutional or economy detail they often do. So before a set drops it is re-checked by models from a different company entirely. A confident cross-family disagreement on a key the first judge let through fails the whole set for the night.
Check 06
No empty slots, and no unchecked filler
A slot whose candidates all die does not get an unchecked question dropped in to fill the gap. It falls back to one that already cleared this same gauntlet on an earlier night and has never been served to anyone.
Where a human sits in this
Every question in every set is readable on an internal review console, with the news item it came from, its answer key and the judge’s verdict beside it. Anything that looks wrong is quarantined in one click, which pulls it out of every future set immediately and keeps its full trail for review.
That is the manual half. The automatic half runs on your feedback. After you answer, every question carries a thumbs up, a thumbs down, a difficulty rating, and a one-tap flag with a reason attached to it (factual error, ambiguous wording, out of syllabus, wrong citation). When enough people mark the same question down, or two flags on it are confirmed as real, a nightly job quarantines it without waiting for me to get to it.
The same job recalibrates difficulty. Once a question has been answered enough times, a hundred graded attempts, its real correct-rate replaces the generator’s guess about how hard it was, and that lands back in how the next night’s questions are aimed.
The pipeline is also watched in aggregate rather than one question at a time. How often the judge disagrees with the generator, cut by subject and by question format. How far realised difficulty has drifted from what was aimed at, and whether the correct answer is creeping toward one position out of four. Movement in those numbers is the early warning that something upstream has started to rot, and it shows up well before any single question looks wrong.
What we do not claim
No AI system is perfect and this one is not either. Every layer above exists because the layer before it can be wrong. A grounded checker can be handed a stale page. A judge can share a misconception with the model that wrote the question. A question can be technically correct and still be a poor question to practise on.
So the honest position is that some errors will reach you. When one does, tap “Something off?” under the question and pick a reason. Flagged questions are reviewed within a day, and a confirmed error is pulled from circulation so nobody sees it again.
What does not happen is a quiet rewrite of history. Answers you have already submitted stay scored the way they were scored, and a finished day’s leaderboard is not recomputed behind everyone’s back. Pulling the question protects everyone who has not seen it yet, which is the group that can still be protected. Rescoring a day you already walked away from would make your own record mean nothing.
Previous year questions are a separate case. They are the real paper, so they are not generated and not judged. They are ingested as published and used as the ground truth everything above is measured against.
Written by Ashish Noel, who sat this exam twice before building this. That part is on the about page. If something here is wrong, or a question deserved a better explanation than it got, write to hello@prelimo.in.