How to feed your AI to keep it from hallucinating

Kiran Prakash·Written with

The way you prepare a file matters more than which AI tool you upload it to. Research from 2024 to 2026 shows that small formatting choices can swing an AI’s accuracy by double digits, and that most “made-up” answers trace back to what was in the file, where it sat, and how the question was asked.

You have probably done this already. You drop your lecture notes, a PDF textbook chapter, or a company handbook into NotebookLM, ChatGPT, Claude or Gemini. Then you ask: “What are the three causes of inflation in chapter 4?” Sometimes the answer is perfect, with a citation. Sometimes it is confident, fluent and wrong.

That wrong answer has a name: a hallucination. It means the AI produced something that sounds right but is not supported by your sources or by reality. This post explains, in plain English, why it happens and what you can change in your own files and questions to make it much rarer. Each technical term is in bold and explained the first time it appears, and there is a glossary at the end.

What actually happens when you upload a file

The AI never sees your file the way you do. It sees a long strip of text pieces, and often only some of them.

The model behind these tools is a large language model (LLM): a program trained on huge amounts of text to predict the next word. Here is roughly what happens between “upload” and “answer”:

  1. Conversion. Your PDF, slide deck or spreadsheet is turned into plain text. This step is called parsing or document conversion. Layout, colours and pictures are mostly thrown away here.
  2. Tokenisation. The text is cut into tokens, small chunks of roughly three-quarters of a word each. Everything the AI reads and writes is counted in tokens.
  3. Chunking. For large uploads, the text is split into short passages called chunks, often a few hundred words each.
  4. Retrieval. When you ask a question, the tool searches your chunks and pulls out the ones that look most relevant. Searching by meaning uses embeddings, which turn text into numbers so similar ideas sit close together. Searching by exact words is called keyword search (a common method is BM25).
  5. Generation. The selected chunks, plus your question, are placed into the model’s context window: its short-term working memory for this conversation. The model writes an answer using only what is in that window and what it learned during training.

Steps 3 to 5 together are called RAG (retrieval-augmented generation). Many “chat with your documents” tools work this way, though each one handles the details differently. The key point is simple: if the right passage is garbled in step 1, cut in half in step 3, or not picked in step 4, the model cannot use it. It will then either say it does not know, or fill the gap with a guess.

Markdown, JSON or plain text: does it matter?

A little, but much less than people think. For today’s big models, clean structure matters more than which syntax you pick.

A few format names you will hear:

  • Plain text (.txt): just words, no formatting.
  • Markdown (.md): plain text with light symbols for structure, such as # for a heading and - for a bullet. Most AI tools read and write it naturally.
  • JSON and YAML: ways of writing data as labelled fields, like name: Priya and age: 34. Programmers use them to store records.
  • XML tags: labels wrapped around content, like <notes>...</notes>, to mark where one thing starts and ends.
  • CSV: a spreadsheet saved as text, with values separated by commas.

What the research found:

  • Formatting can matter a lot for smaller models. A 2024 ICLR paper found that changing only spacing, punctuation and separators moved one model’s accuracy by up to 76 points. The typical swing was about 10 points (Sclar et al.).
  • Bigger models care less. A 2024 study found GPT-3.5 swung by up to 40% depending on the prompt format, while GPT-4 stayed far more consistent (He et al.).
  • There is no single winner. Different benchmarks crowned Markdown key-value lists, YAML and JSON in different tests, and XML ranked second of 11 in one test and last in another (Improving Agents, tables; Improving Agents, nested data; FileConcat review).
  • Labels beat bare grids. In one 2025 test of the same data in 11 formats, writing each record as labelled lines (“Price: $40”) scored 60.7%. A bare CSV scored 44.3% (Improving Agents). When a value sits far from its column header, the model has to work out which column it belongs to, and it sometimes gets that wrong.

The practical takeaway: for notes and documents, use Markdown-style headings and bullets. For lists of facts or records, put a label next to every value. Do not spend time converting everything to JSON.

Everyday formatting rules that work

Write your file so that any single paragraph still makes sense when it is read on its own, because that is often how the AI will see it.

Remember chunking: the tool may hand the model only one passage from page 14, without page 13. So:

  1. Use real headings. Start each topic with a heading, such as ## Photosynthesis: light reactions. Tools that split along headings keep related ideas together, and the heading tells the model what the passage is about.
  2. One topic per section. Do not mix the exam dates into the paragraph about mitochondria.
  3. Name things instead of pointing back. Write “the 2023 Finance Act” rather than “the act mentioned above” or “it”. A chunk that says “it increased by 4%” is useless on its own.
  4. Put a label on every number. Write “Revenue (2024, USD): 1.2 million”, not a lone “1.2” in a grid. Always include units and dates.
  5. State key facts plainly, near the top of the section. “The deadline is 15 March 2027” beats a deadline hidden in the fourth sentence of a story.
  6. Use the words people will ask with. If your notes say “myocardial infarction” but you will ask about “heart attacks”, write both. Research shows models find answers far more reliably when the question and the source share words (more on this below).
  7. Mark dates and versions. Add “Last updated: Sept 2026” to each file, and remove outdated copies. If an old and a new policy disagree, the AI may blend them.
  8. Keep it lean. Remove repeated boilerplate, page footers and near-duplicate copies. Every extra passage is another chance to pull the wrong one.

PDFs, slides and spreadsheets: what gets lost

Plain text survives conversion almost perfectly, but charts and images often do not. If something important lives only in a picture, assume the AI cannot see it.

A 2025 study tested how much information survived when reports were converted for AI use. Text kept over 90% of its information and tables about 90%. Data charts kept only 17% to 70%, and images 8% to 73%, depending on the tool. Slide decks did no better than PDFs (Simmering et al.). A 2026 study of PDF pipelines found that splitting documents along their own headings and attaching labels such as the source, section and date mattered more than which conversion tool was used (From PDF to RAG-Ready).

What you uploadWhat usually goes wrongWhat to do
Typed notes (.txt, .md, .docx)Very littleAdd headings and dates, and you are done
Text-based PDFTwo-column layouts, headers and footnotes get jumbledCheck that the text copies out cleanly. If it does not, paste it into a document instead
Scanned PDF or photo of notesThe tool must use OCR (optical character recognition) to read the image, and it makes mistakes with handwriting and tablesType up the key parts, or run OCR yourself and fix the errors
Slide deckCharts, diagrams and speaker notes are lost or scrambledWrite a sentence under each chart saying what it shows, with the actual numbers
Spreadsheet or CSVRows lose their column meaning, and merged cells breakKeep one header row, give every column a clear name with units, and avoid merged cells
Web page or video linkMenus, ads and comments get mixed in, or only an auto-transcript is usedCopy just the article text, or check the transcript

The single most useful habit: describe every chart in words.For example: “Figure 3: monthly sales rose from 120 units in January to 310 in June, with the biggest jump in April.”

Asking questions the AI can answer well

Upload fewer, more relevant files, and ask specific questions in the same words your sources use. More material is not always better.

Three research findings explain why:

  • Lost in the middle. Models are best at using information at the very start or very end of what they read, and worst at using information buried in the middle (Liu et al., TACL 2024).
  • Context rot. As the amount of text grows, accuracy drops, even on easy tasks. In a 2025 test of 18 models, one conversation-memory task was answered far better from a focused 300-token input than from the full 113,000-token history. Passages that look similar to the answer but are not, called distractors, made things worse (Chroma).
  • Word matching. When a question shares no key words with the passage that answers it, 10 of 12 models tested fell below half their normal accuracy once the text reached about 32,000 tokens. GPT-4o dropped from 99.3% to 69.7% (NoLiMa, 2025).

What to do with that:

  1. Make a focused notebook. One notebook per course, project or topic beats one notebook with everything you own.
  2. Be specific. “What does chapter 4 list as the causes of inflation?” works better than “Tell me about inflation.”
  3. Echo the source’s vocabulary. If the textbook says “cost-push inflation”, use that phrase in your question.
  4. If you paste text into a chat yourself, put the material first and the question last. Anthropic’s guidance for Claude says this can improve answers by up to 30% on long, multi-document inputs (Anthropic docs).
  5. Break big questions into steps. Ask for the relevant facts first, then the comparison or conclusion.
  6. Ask for the answer as normal sentences first. Forcing a strict format, such as a JSON table, while the model is still reasoning can lower accuracy. In one 2024 study, GPT-3.5 scored 76% on maths problems in plain text but 49% when forced into JSON (Tam et al.). Get the reasoning first, then ask for the table.

Why AI makes things up, and how to stop it

AI guesses because it has learned that a confident answer is usually rewarded. Your job is to give it permission not to guess, and to make it show its evidence.

Why guessing happens.OpenAI’s 2025 research argues that models are trained and scored like students on a multiple-choice test. A guess sometimes earns points, while “I don’t know” never does. On one fact-checking test, a model that almost never declined to answer (o4-mini) was wrong 75% of the time. A newer model that declined about half the time was wrong only 26% of the time (OpenAI, Sept 2025).

Why your upload can make it worse. Google researchers found that giving a model passages that are on topic but missing the actual answer can increase hallucinations. In one case, wrong answers jumped from 10% with no documents to 66% with these incomplete ones. The extra text made the model more confident without making it more correct (Google Research, ICLR 2025). The term for having everything needed is sufficient context.

What you can do, as copy-paste phrases:

  • Allow “I don’t know”: “Answer only from my sources. If the answer is not in them, say ‘Not in the sources’ instead of guessing.”
  • Ask for citations: “Cite the source and section for every claim.” Tools like NotebookLM already show citations, so click them and check. This is called grounding: tying each statement to a specific piece of evidence.
  • Quotes first: “First quote the exact sentences that answer this, then give your answer.” Anthropic recommends this for long documents (Anthropic docs).
  • Self-check: “Now list each claim you made and check it against the quotes. Correct anything unsupported.” This is a simplified version of a technique called Chain-of-Verification, which reduced hallucinations in 2024 research (Dhuliawala et al., ACL 2024).
  • Check sufficiency: “Do my sources contain enough information to answer this fully? If not, what is missing?” Ask this before the real question.

Hallucination is also a reason to keep files current. If two versions of a document disagree, the AI may confidently combine them into something neither one says.

Before and after: a page of study notes

The same facts become far more reliable once each one carries its own label and context.

Before (typical quick notes):

Week 5
- it peaked at 11.1 in oct then fell
- 3 causes (see slide)
- prof said this is on the exam!!
- chart on p.12 important

If the AI retrieves only this chunk, it cannot tell what “it” is, what the three causes are, or what the chart shows. Asked “When did UK inflation peak?”, it may guess.

After (AI-ready):

# Macroeconomics 101 – Week 5: Inflation
Source: Lecture 5 slides, Dr Rao. Last updated: 2026-09-28.

## UK inflation rate, 2022
- UK CPI inflation (consumer price inflation) peaked at 11.1% in October 2022, then fell. [Example figure; check your own slides.]

## Three causes of inflation (from slide 8)
1. Demand-pull inflation: too much spending chasing too few goods.
2. Cost-push inflation: rising production costs, such as energy, push prices up.
3. Built-in inflation: wages and prices rise together in a loop.

## Figure on page 12 (described in words)
Line chart of UK CPI, 2019–2024: below 3% until mid-2021, rising to a peak of 11.1% in late 2022, back near 4% by the end of 2023.

## Exam note
The professor said the three causes of inflation will be on the exam.

Every section now makes sense on its own. Numbers have labels, units and dates. The chart is in words. The terms you will search for (“causes of inflation”, “cost-push”) appear in the text.

Cheat sheet and glossary

Before you upload:

  • Typed text, not a photo or scan, wherever possible
  • A heading for each topic, with one topic per section
  • Names instead of “it” or “the above”
  • A label, unit and date on every number
  • Every chart and image described in words, with its numbers
  • A source and “last updated” date at the top, and old versions removed
  • Only the files this question needs

When you ask:

  • Be specific, and use the same words as the source
  • Say “answer only from my sources, or say ‘Not in the sources’”
  • Ask for quotes or citations, then check them
  • Ask for reasoning in plain sentences before asking for a table
TermPlain-English meaning
BM25A classic keyword-search method used to find passages that contain your exact words
Chunk / chunkingA short passage cut from your file / the act of cutting files into those passages
Context windowThe AI's short-term memory for one conversation: everything it can see at once, measured in tokens
Context rotThe drop in accuracy as you give the AI more and more text
CSVA spreadsheet saved as plain text, with values separated by commas
DistractorA passage that looks related to your question but does not answer it
EmbeddingA list of numbers representing a passage's meaning, so similar ideas can be found together
GroundingTying every statement in an answer to a specific source
HallucinationA confident answer that is not supported by the sources or by the facts
JSON / YAMLFormats for writing data as labelled fields, such as name: Priya
LLM (large language model)The AI engine behind ChatGPT, Claude, Gemini and NotebookLM
Lost in the middleThe tendency to overlook information placed in the middle of a long input
MarkdownPlain text with simple symbols for structure, such as # for headings
OCROptical character recognition: turning a picture of text into actual text
ParsingConverting a file such as a PDF into plain text the AI can read
RAG (retrieval-augmented generation)Finding relevant passages in your files, then having the AI answer using them
Sufficient contextYour sources contain everything needed to answer, not just related material
TokenThe unit AI reads in, about three-quarters of a word
XML tagsLabels such as <notes>...</notes> that mark where a piece of content starts and ends

Sources

The format benchmarks from Improving Agents and FileConcat are practitioner tests on small models, not peer-reviewed papers. Exact numbers vary by model and change as models improve.