Context engineering is the practice of curating exactly which tokens Claude sees at inference time — and Agent Skills are the cleanest tool for it, because a Skill keeps only a one-line trigger in context and loads its full instructions on demand. If you've been stuffing everything into one giant system prompt, Skills are the fix: install as many as you like, pay ~100 tokens each while they sit idle, and let the right one inject deep, task-specific expertise the moment your request matches it. Here's the model and the exact syntax.

Why context engineering replaced prompt engineering#

Anthropic frames context engineering as "the natural progression of prompt engineering." Prompt engineering is wording a single instruction well. Context engineering is managing the entire set of tokens — system prompt, tools, retrieved data, and message history — across a whole multi-turn run.

The reason it matters is physical: a model's attention is a finite budget, and "every new token introduced depletes this budget." Worse, there's context rot — as the window fills, "the model's ability to accurately recall information from that context decreases." A 20,000-token system prompt that's 90% irrelevant to the current task doesn't just waste money; it makes Claude worse at the 10% that counts. The goal is always the smallest set of high-signal tokens for the job in front of you.

That's the exact problem Skills solve.

What a Skill is#

An Agent Skill is a folder containing a SKILL.md file — YAML frontmatter plus a markdown body — and optionally some scripts or reference files. The minimum is two required fields and a body:

---
name: pdf-processing
description: Extract text and tables from PDF files, fill forms, merge documents. Use when working with PDFs, forms, or document extraction.
---

# PDF Processing

## Instructions
Use the bundled scripts to extract, fill, or merge PDFs. Prefer
`fill_form.py` for AcroForm fields; fall back to text extraction only
when a form has no fillable fields.

## Examples
- "Pull the tables out of this report" → extract, return as markdown.
- "Fill this application PDF from the JSON" → run fill_form.py.

Two rules make or break a Skill:

The mechanism: progressive disclosure#

Here's why a Skill costs almost nothing until it's needed. Skills load in three levels:

  1. Metadata (always loaded). Only the name + description sit in the system prompt — about ~100 tokens per Skill. Until a Skill triggers, that's its entire footprint.
  2. Instructions (loaded on match). The moment your request matches the description, Claude reads the full SKILL.md body from the filesystem — under 5k tokens.
  3. Resources (loaded as needed). Bundled files and scripts load only when actually used. And when Claude runs a bundled script, only its output enters the context — the script's source never does.

The payoff, in Anthropic's words: "this lightweight approach means you can install many Skills without context penalty." That is context engineering in a box — the just-in-time retrieval tactic, packaged so you don't have to wire it yourself. You get deep expertise injected exactly when relevant, instead of a megaprompt that's mostly dead weight on every call.

Where Skills live, and where they work#

In Claude Code, Skills are filesystem-based — no upload step:

Skills also work on claude.ai and the API (referenced by skill_id, with the code execution tool). One caveat: custom Skills don't auto-sync between surfaces, so you install them where you need them. If you're setting up Claude Code itself, our Claude Code in VS Code setup guide covers the editor side.

The rest of the toolkit: managing the window directly#

Skills handle injecting the right context. For controlling the whole window in a live session, Claude Code gives you four commands:

The mental model: **Skills and just-in-time retrieval decide what comes in; /compact and /clear decide what stays; CLAUDE.md and the memory tool decide what survives.** Get those three flows right and you're doing context engineering, whatever you call it.

The one thing to do today#

Take your longest, most-repeated instruction — the coding-style rules, the report format, the review checklist you paste every time — and move it out of your prompt into a Skill. Give it a sharp description (what + when), drop it in .claude/skills/, and watch /context afterward. You'll see the same expertise, on demand, for ~100 idle tokens instead of thousands on every call. That's the whole game: fewer tokens, higher signal, expertise that shows up exactly when the work needs it.

For the adjacent piece of the puzzle — how models remember across turns and how to read the claims — see our guide on how to read an agent memory benchmark. And once your context is lean, the next lever is cost: route each request to the cheapest capable model.