The questions below come up often in Gen AI and AI engineering rounds, grouped by area. For each, you get what a strong answer covers, not a script. Interviewers follow up on whatever you say, so understand each point well enough to defend it one level deeper.
LLM fundamentals
1. What is a token, and why does tokenisation matter?
A token is the unit the model reads and writes, usually a piece of a word produced by a subword scheme such as byte-pair encoding. Then explain why you care: context limits, latency and cost are counted in tokens, not words, and the same sentence can cost more tokens in some languages and scripts than in English. Because numbers and rare words get split into odd pieces, tokenisation also explains failures like miscounting letters.
2. What does the context window limit, and what happens at the edge?
The context window is the maximum number of tokens the model can handle in one call, and it covers the prompt and the generated output together. Go past it and something gets cut: either the API rejects the request or your code truncates the input. Mention that a bigger window is not free. Long prompts are slower and cost more, and models can be less reliable at using facts buried in the middle of a long context. That is why retrieval still matters when windows are large.
3. What do temperature and top-p do?
Temperature scales the model's output scores before they become probabilities. Low temperature sharpens the distribution towards the most likely tokens, high temperature flattens it. Top-p (nucleus sampling) samples only from the smallest set of tokens whose combined probability reaches p. Good answers add two practical points: use low values for extraction, classification and code, and higher values for brainstorming; and usually tune one of the two, not both. Bonus point: even at temperature zero, outputs are not always perfectly repeatable in practice.
4. Why do models hallucinate?
The model is trained to produce likely text, not to check facts. When it lacks the information (a knowledge cutoff, a training gap, your private data), it still produces fluent text, and nothing in the basic setup makes it say "I don't know". Move straight to mitigations: ground it in retrieved sources, ask for citations, explicitly allow it to decline, check claims against the sources, and measure how often it happens.
Prompting
5. When do you use few-shot examples instead of instructions?
Use examples when the format or judgement is easier to show than describe, such as a particular tone or an edge-case classification rule. Then name the costs: examples use tokens on every call, and the model copies their surface features, so unbalanced or look-alike examples bias the output. Picking diverse examples, or retrieving the most similar examples for each input, is the kind of detail that marks real experience.
6. How do you get reliable JSON or structured output?
- Use the provider's structured-output, JSON mode or function-calling features where available, since these constrain what the model can produce.
- Keep the schema small and use enums for fields with fixed values.
- Validate every response against a schema in code, and on failure retry with the validation error included.
- Know that valid JSON is not the same as correct content. The fields still need checking.
7. What goes in the system prompt, and what does not?
The system prompt holds stable instructions: role, rules, output format, what to do when unsure. User input and retrieved text go in separately and should be treated as data. The point interviewers listen for is that the system prompt is not a security boundary. A determined user or a poisoned document can still override it, so you never put secrets in it or rely on it alone to block actions.
Retrieval-augmented generation (RAG)
8. How do you choose a chunking strategy and chunk size?
Chunks that are too small lose the surrounding context the answer needs. Chunks that are too large blur the embedding across several topics and fill the prompt with irrelevant text. Prefer splitting on document structure (headings, paragraphs, table boundaries) over fixed character counts, add some overlap, and attach metadata like source and section title. Mention small-to-big retrieval: match on small chunks, pass the larger parent section to the model. Finish with: "I would pick the size by measuring retrieval quality."
9. How do embeddings and vector search work?
An embedding model maps text to a vector so that similar meanings sit close together, measured by cosine similarity or dot product. At scale you use approximate nearest neighbour indexes such as HNSW or IVF, through tools like FAISS or pgvector, which trade a little recall for a lot of speed. Two details score well: queries and documents must be embedded with the same model, and changing the embedding model means re-embedding the whole corpus.
10. Why add BM25 and use hybrid search?
Dense retrieval is good at meaning but weak on exact strings: error codes, product IDs, names, acronyms. BM25, a keyword ranking method, catches those. Hybrid search runs both and merges the results, commonly with reciprocal rank fusion, which combines ranks instead of raw scores on different scales. Say which queries made you reach for it.
11. What does a reranker add?
The first-stage retriever is a bi-encoder: it embeds query and document separately, so it is fast but coarse. A reranker, usually a cross-encoder, reads the query and each candidate together and scores relevance far more accurately, but it is too slow to run over the whole corpus. So you retrieve a few dozen candidates cheaply and rerank them down to the handful that go into the prompt. Mention the latency it adds.
12. How do you evaluate retrieval?
Build a labelled set of real questions, each mapped to the chunks or documents that answer it. Then measure:
- Recall@k: how often a relevant chunk appears in the top k. This matters most, because the model cannot use what was never retrieved.
- Precision@k: how much of the top k is actually relevant.
- MRR or nDCG: whether relevant results sit near the top.
The key insight is to evaluate retrieval separately from generation. When an answer is wrong, you need to know whether the right chunk was missing or the model ignored it.
Fine-tuning, RAG or prompting
13. When would you fine-tune instead of using RAG or better prompts?
Start with prompting, because it is the cheapest to change. Use RAG when the problem is knowledge: facts that change, private documents, answers that need citations, or per-user access control. Fine-tune when the problem is behaviour: a consistent format, tone or narrow task that prompting cannot hold reliably, or when you want a smaller, cheaper model to match a larger one on that task. Say clearly that fine-tuning is a poor way to add facts that change, and that it needs good training data, an eval set and a plan for retraining.
14. What is LoRA, and why is PEFT popular?
Parameter-efficient fine-tuning (PEFT) updates a small number of parameters instead of the whole model. LoRA freezes the base weights and learns a pair of small low-rank matrices whose product is added to selected weight matrices, often the attention projections. You train far fewer parameters, need less memory, and get a small adapter file you can swap per task on one base model. QLoRA applies the same idea on top of a quantised base model to save more memory.
Evaluation
15. How do you know your LLM feature is good enough to ship?
This is the question that separates candidates. A strong answer has an eval story:
- A golden set of representative inputs, including edge cases and adversarial ones, with expected outputs or grading criteria. Rerun it on every prompt, model or retrieval change.
- Groundedness or faithfulness: is every claim in the answer supported by the retrieved context? This is different from correctness, and both matter.
- LLM-as-judge, with its biases named: preference for whichever answer comes first, for longer answers, and for text that resembles its own output. Mitigate with clear rubrics, pass/fail criteria, swapping answer order, and checking the judge against human labels.
- Online signals after launch: user feedback, escalations, and sampled conversations reviewed by people.
Production concerns
16. Your RAG app is too slow and too expensive. What do you change?
Separate the two, then measure where time and tokens go before changing anything.
- Latency: stream tokens so the user sees output early; shorten prompts; run retrieval steps in parallel; route easy requests to a smaller model; cap output length.
- Caching: exact-match response caching, prompt or prefix caching where the provider supports it, and semantic caching only with care, since a near-match can return a wrong answer.
- Cost: input and output tokens both cost money. Set a token budget per request, send fewer and tighter chunks, and batch offline jobs.
- Observability: trace each request with the prompt, retrieved chunks, output, latency and token counts, with personal data redacted, so you can debug a bad answer after the fact.
Agents and tool calling
17. How does function calling work, and how do agents fail?
You describe tools with names and argument schemas. The model returns a structured request to call one; your code runs it and sends the result back; this loops until the model gives a final answer. The model never executes anything itself. Then list the failure modes, because that is what the follow-ups target:
- Wrong tool, invented arguments, or calls to tools that do not exist.
- Loops that never finish, and small errors that compound over many steps.
- Context filling up with tool output until the original goal gets lost.
- Side effects that cannot be undone, such as sending, deleting or paying.
Fixes: validate arguments, cap steps and set timeouts, keep the tool list short with clear descriptions, give tools least privilege, and require human confirmation for irreversible actions.
Safety
18. What is prompt injection, and how do you defend against it?
Direct injection is a user typing instructions that override yours. Indirect injection is worse: the instructions sit inside a retrieved document, web page, email or tool result, and the model reads them as if they came from you. Be honest that there is no complete fix, then describe layers of defence:
- Mark retrieved text clearly as data, and never put secrets in the prompt.
- Limit what tools can do, and confirm sensitive actions with the user.
- Block exfiltration paths, such as rendering links or images to arbitrary URLs.
- For personal data: redact before logging or sending to third parties, and apply document permissions at retrieval time so a user can never retrieve what they could not open.
The design question: a RAG chatbot over company documents
Expect some version of "design a chatbot that answers employee questions from internal documents". A good answer names these components and the tradeoff at each:
- Ingestion: connectors, parsing PDFs and tables, chunking, metadata, and the access permissions for each document.
- Indexing: embeddings in a vector store plus a BM25 index, with incremental updates when documents change or are deleted.
- Query path: optional query rewriting, hybrid retrieval filtered by the user's permissions, reranking, then a prompt with numbered sources.
- Generation: answers with citations, and an explicit "I couldn't find this" when retrieval comes back weak.
- Evaluation: a retrieval set and an answer set, both rerun before every change.
- Operations: tracing, feedback buttons, latency and cost budgets, and injection defences for document content.
How answers lose marks
- Buzzwords without tradeoffs. "I'd use RAG with a vector database" says nothing. Say what you chose, what it costs, and when you would choose otherwise.
- No eval story. If your answer to "how do you know it works?" is "I tested a few queries", the interviewer stops trusting the rest of the design.
- Ignoring cost and latency. Every extra model call, reranker or longer prompt has a price in rupees and seconds. Strong candidates mention it before being asked.
- Treating the model as reliable. Designs that assume valid JSON, correct tool arguments or harmless retrieved text break in the first follow-up.
How to practise this
These topics get tested through follow-ups: mention reranking, and the next question is what it costs; mention LLM-as-judge, and the next is how you know the judge is right. Practise by saying your answers out loud to someone who keeps asking "why".
On mangoose.tech you can sit a live mock interview for Gen AI Engineering or AI Engineering roles. You speak your answers to Meera, who asks follow-ups and keeps pressing one level past each answer. Afterwards, a written gap report marks where your answer broke down, such as a tradeoff you never mentioned or a claim you could not defend.
Practise a Gen AI engineering round out loud
Answer LLM and RAG questions with Meera, handle her follow-ups, and read a report showing exactly where your answers broke down.