AI flashcards: why most of them are worse than none, and the one check that tells you

Kao nguyễn9 min read

The pitch is simple and it works on everyone who has ever tried to write flashcards by hand. Paste a PDF, or a lecture, or a chapter, and get fifty cards back in the time it used to take to write two. The card-writing was the step that killed the habit; now it is gone.

Then, a month later, you are reviewing a card that says the treaty was signed in 1648 and you have no idea whether that is what the book said or what the generator decided sounded right. Or you are forty cards into a session that is mostly definitions of terms you never cared about. Or you have stopped, because there are six hundred due and no one wrote them, so no one is attached to them.

Generated cards are not bad because they are generated. They are bad in three specific ways, all three come from the same design choice, and there is one check that catches all of them before a card gets into your queue. This is the three ways, the check, and how to get cards worth keeping out of a general model like ChatGPT when a purpose-built tool is not what you want.

The three ways they go wrong

Too many

A generator that is paid per card, or judged by a demo that shows a big number, will make as many as it can. Fifty from a chapter. Four hundred from a book. The number looks like value and it is the opposite: the twelve cards that mattered are now buried in the three hundred and eighty that did not, the daily session is dominated by things you did not want, and the backlog that ends most spaced repetition habits arrives in the first week instead of the third.

The tell is that the count does not change with the source. A dense chapter and a thin one both come back as fifty. A generator that was actually reading would produce a dozen from one and three from the other.

Recognition-shaped

The easiest card to generate is a definition. "What is X?" with the sentence that defined X on the back. It is easy because the source has one sentence per term and the term is in the sentence, so the generator can produce it without understanding anything.

It is also close to useless, because it asks for recognition. A reader who saw the term once can answer "what is X" by reconstructing it from the word. The card that would have kept the idea asks for the mechanism, the reason, the contrast: not "what is the testing effect" but "why did the group that was tested once remember more a week later than the group that reread three times". Active recall needs a question that cannot be answered from the cue alone, and generators, left to themselves, do not write those.

Unverifiable

This is the one that does harm rather than merely wasting time.

A card says something. You review it, you recall it, the schedule pushes it out to next month, and by the third review it is simply a thing you know. If the card was wrong, if the generator misread a number, conflated two paragraphs, or filled a gap with something plausible, you now know a wrong thing with the confidence of someone who has reviewed it four times.

There is no way to catch this at review, because at review you are testing yourself against the card, not the card against the source. The only moment to catch it is before the card enters the queue, and to catch it then you need to see the passage it came from. A card without its passage cannot be checked, and a card that cannot be checked is indistinguishable from an invented one.

The complaint people make about generated cards on forums, that they are inefficient and that content gets lost or garbled, is this failure mode described from the outside.

The one check

Can you click from the card to the sentence it was written from?

That is the whole test, and it catches all three failures at once. A card that cites its passage can be checked for accuracy, so the unverifiable problem goes away. A card that cites a passage which turns out to be a definition is visibly a definition card, and you drop it. And a generator that has to cite a passage for every card cannot produce fifty from a chapter that has twelve things in it, because there are only twelve passages worth citing.

The check also tells you something about the generator's design. One that works from the topic, or from a summary of the source, or from the source in pieces it later throws away, cannot cite anything, because it never held the passage and the card together. One that reads the source and writes each card with the sentence beside it can. This is a structural difference and it shows up in whether the citation exists.

As of September 2026, most generators do not offer this. The tool we make, Agent Memo, was built around it: every card the agent proposes carries the passage it came from, the proposals are shown to you before any enter review, and dropping one is a click. Weigh that paragraph knowing whose it is. The check is the point, and it applies to any tool, including ours.

A closed book connected by teal threads to three blank cards.

The source-to-card link is what makes a generated claim checkable.

Using ChatGPT for flashcards without the three failures

If you would rather use a general model than a purpose-built tool, you can get most of the way with the prompt. The defaults produce all three failures; the corrections are simple and they are all about giving the model less room.

Paste the passage, not the topic. "Make flashcards about the French Revolution" gets you cards from the model's memory of the topic, which is exactly the unverifiable case. Paste the two pages you actually read and ask for cards from those pages only.

Cap the count and make it earn each card. "At most eight cards. Only for something a careful reader would want to still know in a year. Skip definitions of terms." The cap forces selection, and the selection is where the value was.

Require the citation. "For each card, quote the sentence from the passage it is based on, verbatim." Now every card can be checked, and a card whose quoted sentence does not support it is visibly wrong. This one instruction removes the failure that does harm.

Ask for mechanism, not labels. "Each question must require explaining why or how, not naming what." This is the difference between recognition and recall, and models will follow it if told.

Review before you import. Read every card against its quoted sentence. Drop the ones that are definitions, the ones you do not care about, and the ones whose sentence does not say what the card says. Ten minutes here saves a year of reviewing something wrong.

A prompt that does all of this, for a pasted passage:

From the passage below and nothing else, write at most eight flashcards for a reader who wants to still know this in a year. Skip definitions. Each question must require explaining why or how. For each card, quote verbatim the sentence it is based on. Format: Q, A, Source.

You will get fewer cards than the default, and they will be the ones worth having.

Three approaches, compared

Paste into a general model A flashcard generator Agent Memo
Reads the whole source Only what you paste, in pieces Often a chunk or a summary Yes, the whole thing
Cites the passage per card Only if you ask Rarely Every card
Caps the count Only if you ask Usually not; more is the demo A dozen from a book; a ceiling per capture
You review before it enters the queue You have to do it yourself Sometimes Always; dropping is a click
Who it is for Someone who will prompt carefully every time Someone with an exam and a PDF Someone who reads and will not write cards

The middle column describes a category rather than any one product, and the category is changing; check the specific tool against the one test rather than against this table.

What a good generated card looks like

From a passage explaining that the Meiji government abolished the han domains in 1871 and replaced roughly 260 of them with prefectures under centrally appointed governors, so that tax and conscription flowed to Tokyo rather than to the old lords:

Q: What replaced the han domains in 1871, and what did the central government gain by doing it?

A: Prefectures under governors appointed from Tokyo. Tax and conscription now went to the centre rather than to some 260 local lords.

Source: "…replaced roughly 260 of them with prefectures run by centrally appointed governors, so that tax and conscription flowed to Tokyo instead of to the old lords."

It asks for the mechanism. It carries the number, and you can see where the number came from. If the book had said 270, the source line would say 270, and you would catch the card. That is the whole difference between a card you can trust and one you have to.

Frequently asked questions

Are AI-generated flashcards worth using at all? Yes, if they pass the check: you can see the passage each one came from, and you reviewed them before they entered your queue. Without that, they are a fast way to memorise a generator's mistakes.

Can ChatGPT make good flashcards? It can, from a pasted passage, with a capped count, a citation requirement and an instruction to ask for mechanism rather than labels. Its defaults produce the three failures above; the prompt is what changes that.

How many cards should a generator produce from a chapter? As many as the chapter has things worth keeping, which is usually far fewer than it produces. A dozen from a serious non-fiction book is a good yield; fifty from a chapter is a sign nothing was selected.

What is the difference between an AI flashcard generator and Agent Memo? Agent Memo is a generator that reads the whole source, proposes a small set, cites the passage for every card, and shows you all of them before any enter review. Most generators do some of these. The citation is the one that lets you check the others. It is also the tool we make, so verify that description against the product.

Will generated cards be worse than ones I write myself? A card you wrote from a passage you understood will usually be better than a generated one, and it will usually not exist, because the writing is the step that fails. A generated card that cites its passage and survived your review is the realistic alternative to no card at all.

What to do now

Take the last thing a generator or a model produced for you and, for each card, try to find the sentence it came from. The cards where you cannot are the ones to delete. Then, next time, ask for the sentence up front, or use a tool that gives it to you without asking, and spend the review on the cards that earned their place.