Most people prompt: they open a Claude or ChatGPT tab, describe what they need that hour, and get the same quality of answer anyone else with a subscription can get. Engineering a custom Claude means building the structure around the conversation — a system prompt that sets the role and the rules, a curated knowledge base, Skills for repeatable work, tools and connectors into company systems, an evaluation set that defines what good looks like, and approval rules for anything with consequences. Prompting produces a demo. The engineered version produces something a team can rely on at 4pm on a Tuesday without anyone remembering the magic words.

The model underneath is identical in both cases. When two companies get wildly different value from the same Claude, the difference is almost never the model.


What Ad-Hoc Prompting Actually Gets You

Prompting works, and a capable person with a good prompt can get excellent output from Claude on a one-off task. The trouble starts when that output needs to be repeatable, shared, or trusted by someone who wasn't in the chat.

Anthropic's own guidance puts it neatly: treat Claude as "a brilliant but new employee" who lacks context on your norms and workflows.[1] Every ad-hoc chat re-onboards that employee from scratch. The brand voice lives in one marketer's saved prompt, the pricing rules in a sales lead's head, the approved-claims list in a PDF nobody remembers to attach. The answer depends on who is typing and what they thought to paste in. Nothing compounds.


The Engineered Stack, Layer by Layer

An engineered Claude moves that context out of individual heads and into structure the assistant carries into every conversation. The components are documented by Anthropic; the craft is in what goes into them.

System Prompts and Project Knowledge

The system prompt is where the role, the audience, the tone and the hard rules live. Anthropic notes that setting a role in the system prompt focuses Claude's behaviour and tone, and that explaining the reason behind an instruction helps Claude generalise from it.[1] The best system prompts read like a thorough onboarding document for a senior hire.

For teams working in Claude's apps, Projects give each use case its own workspace with its own chat history, knowledge base and instructions.[2] On paid plans, Claude switches to retrieval when project knowledge approaches the context limit, which Anthropic says expands capacity by up to ten times.[2] Capacity is not the hard part, though. Curation is: a knowledge base stuffed with every superseded policy will answer confidently from the wrong version.

Skills for Repeatable Work

Skills package instructions, metadata and optional resources such as scripts and templates into a folder that Claude uses automatically when a task calls for it.[3] Only a short description of each Skill sits in context until the Skill is needed, at which point the full instructions load; Anthropic calls this progressive disclosure.[3] Skills work across the Claude apps, Claude Code and the API.[3]

In practice, a Skill is where the house style or the proposal structure goes, so the fifteenth document comes out like the first. Anthropic is blunt that Skills should only come from trusted sources, because a malicious one can direct Claude to use tools or run code in ways that don't match its stated purpose.[3] Installing one deserves the care you would give installing software.

Tools and MCP Connectors

Tool use is what lets Claude act rather than only answer. Claude decides when to call a tool based on the request and the tool's description, then returns a structured call that your application executes.[4] That makes tool descriptions part of the prompt, and Anthropic's engineering team argues they deserve as much prompt-engineering attention as the main prompt does.[5]

The Model Context Protocol, which Anthropic released in November 2024 as an open standard for connecting AI assistants to the systems where data lives, is how most of those connections now get built.[6] Connectors in Claude inherit each person's permissions from the connected service, and Team and Enterprise owners can limit which actions a connected service may take.[7] Custom connectors can point at services Anthropic hasn't verified, and the help centre warns that a malicious MCP server can carry hidden instructions.[8] On the API, the MCP connector (currently in beta) lets a build allowlist specific tools and switch the rest off.[9] That is our default: expose the tools the job needs and nothing more.


Evaluation Sets: Where Prompting Becomes Engineering

This is the layer ad-hoc users almost never build. Anthropic's prompt-engineering overview says that before you start you need clear success criteria and a way to test against them empirically, and that if you lack either you should establish them first.[10] It also notes that not every failing test is a prompt problem; latency and cost are sometimes easier to fix by choosing a different model.[10]

An evaluation set is a collection of real inputs paired with a definition of a good answer. Anthropic's guidance is to mirror the real distribution of tasks, include edge cases, and favour more test questions with automated grading over a handful graded by hand.[11] The richest source is discovery itself: the questions asked five times a week, the requests that have gone wrong before, the phrasing customers actually use.

Without an eval set, every prompt change is a guess and every model upgrade a leap of faith. With one, you can change the prompt on Monday morning and know by lunch whether anything broke.


Approvals and Guardrails

Once Claude can act inside company systems, the question moves from whether the answer is good to what the assistant is allowed to do. Remote MCP connectors can read, create, modify or delete data, depending on the permissions granted.[8] Anthropic's advice on agents is to find the simplest solution that works, build in checkpoints where the system pauses for human feedback, and test extensively in sandboxed environments with appropriate guardrails.[5]

We design approvals by consequence. Reading and summarising run freely. Drafting is fine, because a human sees the draft first. Anything that sends, publishes, deletes or commits money goes through a person, every time.

Guardrails on answers matter as much. Anthropic's documented techniques for reducing hallucinations include permitting Claude to say it doesn't know, restricting it to the provided documents, and having it cite quotes for its claims so responses can be audited.[12] Those instructions belong in the system prompt, and the eval set should confirm they hold.


Data Handling Is Part of the Build

Where staff use Claude matters as much as how. Anthropic states that by default it does not use inputs or outputs from its commercial products, including Claude for Work and the API, to train its models.[13] Consumer plans are different: under the terms update Anthropic announced in August 2025, Free, Pro and Max users choose whether their chats are used to improve models, with retention extended to five years for those who allow it.[14] A team doing client work in personal accounts is making a data decision, whether or not anyone signed it off.

On Claude Enterprise, audit logs record events such as sign-ins, project changes and file uploads, and owners can export the previous 180 days.[15] We scope data handling during architecture and review it against the client's AI policy. If that policy doesn't exist yet, write it before the build; our AI policy development work exists for exactly that.


How Mooning Ships Them

Our custom Claude builds follow the same shape each time.

Discovery first. We sit with the people who will use it and ask where time gets lost, which questions come up five times a week, which handovers break and which approvals slow everything down. Then we map the use cases that earn a build and the integrations that need to exist first. A clear brief gets written before a single prompt does.

Then the invisible part. System prompts that hold up under pressure, the tools Claude needs to take real action, a properly indexed knowledge base, approvals wired in where they matter, and logging from day one.

Then the part people see. Internal tools land or stall on whether they feel like part of the team, so every Claude we build gets a name, a voice and an interface that belongs in the client's stack. In our experience the model doesn't decide whether a team uses it. The way it is presented does.

A staged rollout. A pilot group first, then feedback, tuning and expansion, with internal champions equipped and training kept short.

Ongoing tuning. Usage shifts and edge cases surface. Quarterly tune-ups are standard for retainer clients, with bigger revisions when the underlying model takes a genuine step forward. A usable v1 typically takes four to ten weeks from kickoff, depending on how deep the integrations run.


The Same Model, a Different Result

Anyone can buy access to Claude. What nobody can buy off the shelf is an assistant that knows their products, writes in their voice, touches only the systems it should, asks before it acts and gets better the longer it is used. That part is engineering, and it is where the advantage sits.

If your team is still prompting its way through the week, talk to us about what the engineered version would look like.


Frequently asked questions

What is the difference between prompting Claude and building a custom Claude?

Prompting means typing requests into a general Claude account and relying on whoever is typing to supply the context. A custom Claude builds that context into the system: a system prompt with the role and rules, a curated knowledge base, Skills for repeatable tasks, connectors into company systems, an evaluation set and approval rules. The model underneath is the same, but the engineered version gives consistent answers no matter who is asking.

What are Claude Skills and how are they different from Claude Projects?

A Project is a workspace with its own chat history, knowledge base and instructions. A Skill is a packaged folder of instructions, plus optional scripts or templates, that Claude loads automatically when a task calls for it, and it works across the Claude apps, Claude Code and the API. Projects hold context for a use case; Skills hold repeatable ways of doing a task.

Can Claude connect to internal systems like Google Drive or Slack?

Yes. Claude supports connectors built on the Model Context Protocol, an open standard Anthropic released in November 2024, and connectors inherit each user's existing permissions in the source system. Custom connectors can reach internal services as well, but they should only point at servers you trust, and anything that writes, sends or deletes should sit behind human approval.

Does Anthropic train its models on our company's data?

Anthropic states that by default it does not use inputs or outputs from its commercial products, including Claude for Work and the API, to train its models. Consumer Free, Pro and Max accounts work differently, with each user choosing whether their chats are used to improve models. That is one reason company and client work belongs in a commercial account rather than personal ones.

How long does it take to build a custom Claude?

At Mooning, a usable v1 typically takes four to ten weeks from kickoff, faster when scope is tight and longer when integrations run deep. Discovery and a written brief come first, then the build, a staged pilot rollout and ongoing tuning as usage grows.


References