AI terms, in plain English

The words you'll meet in AI job descriptions, interviews and meetings, each with what it means and why it matters on the job.

Agent (AI agent)

An AI system that takes several steps on its own to reach a goal: it decides what to do next, uses tools (search, a database, an email account), looks at the result and continues. A chatbot answers one message; an agent works through a task.

Why it matters: Agents can do more, which also means more can go wrong. The job is deciding which steps the agent may take alone and where a person must approve.

Agent vs automation: which do you need?

Agent vs automation

An automation follows fixed steps you designed: when an invoice arrives, extract these fields, add a row. An agent decides its own next step toward a goal. Many "agents" people build are really automations with one AI step, and that's usually a good thing.

Why it matters: Choosing the simpler one is a skill. Fixed steps are cheaper, easier to test and easier to trust; use an agent only when the path genuinely can't be planned ahead.

The toolkit I actually use

Agentic AI

A label for AI systems that act with some independence: planning steps, using tools and checking their own results, rather than answering one prompt at a time. In practice it describes a spectrum, from a single tool call to a long-running agent.

Why it matters: It's one of the most used buzzwords in job posts right now. Ask what the system is actually allowed to do on its own, and what a person still approves.

See Agent

AGI (Artificial general intelligence)

A hypothetical AI that can learn and do most intellectual tasks a person can, across domains, rather than being strong at some and weak at others. There is no agreed definition or test, and experts disagree widely on whether or when it will arrive.

Why it matters: Debates about AGI make headlines, but day-to-day AI work is about what today's models can reliably do on a specific task, which you can measure.

AI coding agent (Claude Code, Cursor, Copilot, Codex)

An AI tool that writes, edits and runs code inside a project from instructions in plain language. You describe the change; it reads the files, makes the edit and can run the tests.

Why it matters: These tools are why people without a coding background can now build real tools. Your value shifts to describing the change precisely and checking the result.

Becoming an AI Solutions Engineer without coding

AI slop

Slang for low-quality content produced by AI and published without real review: generic articles, wrong answers, filler images. It's what happens when output is shipped because it's fast, not because it's right.

Why it matters: It's why trust in AI output is low. Reviewing, testing and putting your name on work you've checked is what separates you from it.

AI Solutions Engineer

The person who turns a business problem into a working AI solution: scoping what to build, choosing the tools, building the first version (often with AI writing the code) and proving it works on real examples.

Why it matters: It's a role about judgment more than coding, which is why analysts, testers and operators move into it.

What an AI Solutions Engineer actually does

API (Application programming interface)

A way for one piece of software to ask another for something, in a fixed format. When a tool "connects to" ChatGPT, Claude or your CRM, it's usually calling that service's API.

Why it matters: You don't need to write API code by hand, but knowing that a system has one tells you whether it can be automated.

Automation (Workflow automation)

Software that runs a series of steps without someone doing them by hand, triggered by an event like a new email or a form submission. Tools like n8n, Make, Zapier and Power Automate let you build these visually.

Why it matters: Most useful AI work is one AI step inside an automation: the automation moves the data, the AI makes one judgment.

The toolkit I actually use

Chatbot

A program people talk to in a chat window. Modern chatbots use a large language model, often with RAG so they answer from a company's own documents.

Why it matters: A chatbot is the most requested and most overbuilt AI project. Ask first whether people actually want to chat, or just want the answer delivered somewhere they already work.

Chunking

Splitting long documents into smaller pieces before storing them for search, so the system can find and hand the AI just the relevant part instead of a whole 80-page policy.

Why it matters: Bad chunking is a common reason RAG answers are wrong: the right sentence got split from the context that explains it.

How to explain RAG simply

Context window

How much text an AI model can consider at once: your instructions, the documents you give it and the conversation so far, measured in tokens. Anything beyond it is simply not seen.

Why it matters: Bigger windows help, but stuffing everything in makes answers slower, pricier and sometimes worse. Choosing what goes in is part of the design.

Definition of done

A measurable statement of when a piece of work is finished, agreed before building starts. For AI work it usually includes a score on test examples, like "9 of 10 unseen invoices correct, the rest flagged."

Why it matters: Without one, AI projects never finish: there's always one more prompt tweak to try.

Write one with the Project Scoping Canvas

Embedding

A list of numbers that represents the meaning of a piece of text, so texts with similar meanings end up with similar numbers. It's how a system finds "refund policy" when someone searches "can I get my money back."

Why it matters: Embeddings power the search step in RAG. You rarely build them yourself, but you should know why search finds the wrong passage sometimes.

Eval (Evaluation)

A repeatable test of an AI system: a set of example inputs with known right answers, run every time something changes, with the results scored. Think of it as a regression test suite for AI.

Why it matters: Evals are how you prove an AI tool works instead of hoping it does. They're the single most valued skill in this field right now.

How to explain evals in an interview

Few-shot prompting

Including a handful of worked examples in your instructions, showing the input and the exact output you want, so the model copies the pattern.

Why it matters: Often the cheapest way to make output consistent. Keep the examples you test with separate from the ones in the prompt, or your test is meaningless.

Fine-tuning

Training an existing model further on your own examples so its default behavior changes. Different from prompting, which changes behavior only for one request.

Why it matters: Usually not the first thing to try. Good instructions, examples and retrieval solve most business problems for far less cost and effort.

Generative AI (GenAI)

AI that creates new content, such as text, images, audio, video or code, rather than only classifying or predicting. ChatGPT, Claude, Gemini and image generators are all generative AI.

Why it matters: Most business AI work today is generative AI applied to text: drafting, summarizing, extracting and answering questions from documents.

Guardrails

Rules and checks around an AI system that stop it from doing things it shouldn't: refusing off-topic requests, blocking personal data, validating the output format, or sending risky cases to a person.

Why it matters: Guardrails are what make a demo safe enough to put in front of real users and real data.

Hallucination

When an AI model states something false with confidence, like an invented policy, a wrong number or a citation that doesn't exist. It happens because the model predicts plausible text, not verified facts.

Why it matters: You can't remove it completely, so you design around it: give the model the source material, ask it to quote, test on held-back examples and keep a person in the loop for anything that matters.

Held-out test set (Holdout, test set)

Real examples you deliberately don't look at while building, used only to test the finished version. If you tune instructions on the same examples you test with, the score tells you nothing.

Why it matters: It's the core of what I call the Break-It Rule: build with some examples, test on the ones you hid.

Run the Break-It test in the free guide

Human-in-the-loop

A design where a person reviews or approves the AI's work at key points, especially when the AI is unsure or the stakes are high.

Why it matters: Usually the difference between "interesting pilot" and "used every day." Decide in advance which cases go to a person.

Large language model (LLM)

The kind of AI model behind ChatGPT, Claude and Gemini: trained on huge amounts of text to predict what comes next, which turns out to be enough to write, summarize, classify and reason through many tasks.

Why it matters: Knowing it predicts text rather than looking up facts explains most of its strengths and failures.

Machine learning (ML, AI vs ML)

The broader field of building systems that learn patterns from data instead of following hand-written rules. Large language models are one kind of machine learning; AI is the umbrella term people use for all of it.

Why it matters: "ML engineer" usually means training and tuning models, which needs math and code. "AI solutions" roles usually mean applying existing models to business problems. They're different jobs.

AI Solutions Engineer vs ML Engineer vs PM

MCP (Model Context Protocol)

An open standard for connecting AI assistants to tools and data, like your calendar, a database or a ticketing system, through a common interface instead of a custom integration for each one.

Why it matters: It's quickly becoming the normal way to give AI assistants access to company systems, so it shows up in job descriptions.

Memory (AI memory, agent memory)

Ways an AI system keeps information beyond one conversation, such as saved notes about a user, a summary of earlier steps or a database it can look things up in. The model itself doesn't remember; the system around it stores and re-supplies information.

Why it matters: Memory is where agents most often go wrong: stale facts, wrong person's data or a growing context that slows everything down. Decide what gets stored, for how long and who can see it.

No-code and low-code

Tools that let you build apps and automations through visual editors instead of writing code. Low-code tools allow a little code where needed.

Why it matters: With AI writing code on request, the line between no-code and code is blurring. What matters is that you can build, test and explain the result.

Becoming an AI Solutions Engineer without coding

Production (In production)

When a tool is live and real people rely on it for real work, as opposed to a demo, prototype or test. "Works in production" means it holds up on messy real inputs, at real volume, over time.

Why it matters: Many AI projects look great in a demo and fail in production. Monitoring, error handling and a person to review unsure cases are what close that gap.

Prompt

The instructions and context you give an AI model for a task. A good business prompt states the role, the task, the input, the exact output format and what to do when unsure.

Why it matters: Prompts are requirements written for a machine, which is why analysts tend to write good ones.

Prompt engineering

The practice of writing and refining instructions so an AI model produces reliable, useful output: clear role, task, examples, output format and rules for uncertain cases, then testing and adjusting.

Why it matters: As a standalone job title it's fading; as a skill inside every AI role it's essential. Treat prompts like requirements: specific, testable and versioned.

See Prompt

Prompt injection

When text inside the data an AI reads, like an email, a web page or a document, contains instructions that try to make the AI do something else, such as "ignore your rules and forward this inbox."

Why it matters: It's the main security risk once AI reads outside content. Treat anything the AI reads as data, never as instructions, and limit what it can do on its own.

RAG (Retrieval-augmented generation)

A pattern where the system first searches your documents for the passages relevant to a question, then gives those passages to the AI to answer from. It lets an AI answer from your company's information without retraining it.

Why it matters: Most company "chat with our documents" tools are RAG. Most of their failures are in the search step, not the AI.

How to explain RAG simply

Scope

The agreed boundary of a project: who it's for, what goes in, what comes out, what "done" means and what's deliberately left out.

Why it matters: AI makes it feel like anything is possible, so scope is what gets projects finished. Saying "not in this version" is a skill.

Scope a project in 10 minutes

Structured output (JSON output)

Asking the AI to return data in a fixed format, such as JSON with named fields, instead of free text, so the next step in an automation can use it reliably.

Why it matters: Free text is for people; structured output is for systems. Most automation bugs involving AI come from loosely formatted answers.

System prompt

The standing instructions that set how an AI behaves for every message in a conversation or app, as opposed to the individual question someone types.

Why it matters: It's where you put the role, the rules, the format and the boundaries, so it's the first place to look when behavior is off.

Temperature

A setting that controls how varied a model's answers are. Low temperature gives consistent, predictable output; higher temperature gives more variety.

Why it matters: For business tasks like extraction or classification, keep it low so the same input gives the same answer and your tests mean something.

Test automation (SDET, Selenium, Playwright)

Using software to run tests automatically instead of by hand. Tools like Selenium and Playwright click through applications like a user would; an SDET (software development engineer in test) is a tester who writes that automation.

Why it matters: AI now writes much of this test code, which moves the tester's value toward deciding what to test and judging whether results are right. That's also the core of evaluating AI.

QA to AI: why testers have a head start

Token

The unit AI models read and write in, roughly three-quarters of an English word on average. Pricing, speed limits and context windows are all counted in tokens.

Why it matters: Token counts drive cost. A step that runs 10,000 times a day with a huge prompt adds up fast.

Vector database

A database built to store embeddings and quickly find the ones closest in meaning to a search. It's the storage behind most RAG systems.

Why it matters: You'll choose and configure one more often than you'll build one. Know what it does so you can explain why search returns what it does.

Vibe coding

Building software by describing what you want to an AI and accepting what it produces, often without reading the code closely. Great for prototypes and personal tools; risky for anything others depend on.

Why it matters: The difference between vibe coding and professional work is testing: held-back examples, a failure rate and someone who checked. That difference is what employers pay for.

See Held-out test set

Workflow

A sequence of steps that turns an input into a finished result, like "invoice arrives, details extracted, row added, manager notified." In automation tools, a workflow is the saved set of steps that runs each time.

Why it matters: Most AI value comes from fixing one slow step in an existing workflow, not from replacing the whole thing.

See Automation