Custom AI agents and MCP servers that do real work, safely

An agent is a model, a set of tools, and rules about when it may use them. The model is the easy part. We design the tools, the permissions and the approval steps — the same way we built Buddy, which makes real Google Ads changes behind a preview, confirm, execute and receipt loop.

TL;DR

Who this is for

Teams with a repetitive, rules-heavy task that people do today in software you already own, and a clear way to tell whether it was done right. Good first agent projects:

Poor first projects: anything that spends money, emails customers or deletes data without a person approving it, and anything described only as "an AI strategy." If a fixed script or a scheduled workflow would do the job, we'll recommend that instead; Anthropic's own guidance on building agents makes the same point about starting simple.

What's included in a pilot

How we build an agent

One workflow, one metric

Pick a task people do today, measure how long it takes and how often it goes wrong, and define what "working" means before writing code.

Design narrow tools

Separate read tools from write tools. Each tool does one thing with validated inputs — "add negative keywords to this campaign," not "run any query."

Scope the credentials

The agent gets its own credentials with the minimum access the workflow needs, handled in code rather than passed to the model. User-supplied keys are encrypted at rest, as they are in Buddy.

Guard every write

Dry-run by default, a readable preview of what will change, explicit approval, execution, and a receipt in an audit log. Spend caps and rate limits on anything that costs money.

Build an eval set

Real tasks with expected outcomes, run on every prompt or model change, plus adversarial cases that try to make the agent misuse its tools.

Deploy and operate

Usually on Cloudflare Workers or your existing infrastructure, with logs of every tool call reviewed during the first weeks before tools are added.

Agent mistakes we see

What we mean by an AI agent, and where MCP fits

A chatbot answers questions. An agent takes steps: it reads data, decides what to do next, calls a tool, checks the result, and repeats until the job is done or it needs a human. That loop of reasoning and acting is the pattern described in the ReAct paper, and it needs four things beyond the model: tools with clear inputs and outputs, credentials scoped to what the job needs, context about your business, and guardrails on anything that changes data or talks to customers.

The Model Context Protocol (MCP) is the open standard for the tools part. It defines how AI applications connect to external systems, and the project lists support in Claude, ChatGPT, Visual Studio Code, Cursor and others. Build an MCP server for your system once and it works in each of them, which is why most of our tool work ships as MCP servers. Remote MCP servers authenticate users through the OAuth-based flow in the MCP authorization specification, so the agent never handles a raw password.

Security: prompt injection and excessive agency

OWASP lists prompt injection first in its Top 10 for LLM applications. Direct injection comes from what a user types; indirect injection comes from content the agent reads — a web page, a file, an email — that contains instructions. Researchers demonstrated the indirect form against real LLM-integrated apps in 2023, and OWASP is candid that it's unclear whether fool-proof prevention exists. The related risk OWASP calls excessive agency has three root causes: excessive functionality, excessive permissions and excessive autonomy. Our design rules map to those directly.

Risk What we do about it
Excessive functionality Narrow, single-purpose tools; no generic "run SQL" or "call any URL" tool
Excessive permissions The agent's own least-privilege credentials, held in code; read and write split into separate tools
Excessive autonomy Human approval for anything irreversible, costly or customer-facing; spend caps and rate limits
Indirect prompt injection Untrusted content clearly separated and treated as data; an agent that reads it never holds write credentials in the same context
Malformed or unsafe output Output formats validated in deterministic code before any tool runs
MCP-specific attacks No token passthrough, per-client consent, and the other mitigations in MCP's security best practices
Silent regressions Adversarial eval cases run on every prompt or model change

A useful test from Simon Willison's writing: if one agent has access to private data, reads untrusted content and can send data out, it can be tricked into leaking. We break at least one of those three links in every design.

Model choice

We're model-agnostic. Our open-source Python agent supports Claude, GPT and Gemini, and Buddy lets users bring their own keys. We choose per task on tool-use reliability, latency and cost, and the eval set makes switching models a measured decision rather than a guess. Repeated trials matter: benchmarks such as τ-bench show agents that pass a task once can fail it on the next attempt, so we report pass rates across runs, not a single demo.

Pricing and engagement

A short discovery to pick the workflow and define success, a fixed-scope pilot, then production hardening. After launch, either handoff or an ongoing retainer, as with our other app work. Model and hosting costs are billed by the providers to your accounts. Client agents and MCP servers are your IP; we open-source components only when you choose to. Nothing is billed as a percentage of spend; the terms are on our pricing page.

Platforms we work in

Models, protocols and infrastructure we build agents on.

Anthropic logoAnthropic Claude API OpenAI logoOpenAI API Google Cloud logoGoogle Gemini on Google Cloud Model Context Protocol logoModel Context Protocol Cloudflare Workers logoCloudflare Workers TypeScript logoTypeScript GitHub logoGitHub n8n logon8n

What we've built and open-sourced

Production agent

Buddy

Our conversational Google Ads agent: 231 API actions (93 read queries, 138 write mutations) behind a 55-tool catalog, with every write running through preview, confirm, execute and receipt. It runs on the web and in our iOS and Android apps.

How it's built →
Open source · MCP

google-ads-mcp

A Python MCP server giving any MCP client live Google Ads API access. At launch it had 29 tools — 9 read, 7 audit, 11 write, 2 documentation — with every write tool dry-run by default.

View on GitHub →
Open source · Gemini CLI

google-ads-gemini-extension

A Gemini CLI extension with MCP tools, custom commands, agent skills, credentials stored in the system keychain, and hooks that block unconfirmed writes and keep an audit trail.

View on GitHub →
Open source · Agent skills

Claude and Gemini skills

77 reusable agent skills, 2 MCP servers and 7 commands for managing Google Ads from Claude and Gemini CLI, published under the itallstartedwithaidea organization.

Read the overview →
Agent in operations

Squeaky Clean Turf migration

An agent with a browser bridge edited CRM workflows in a builder with no public API, then verified fourteen published workflows against their execution logs before anyone declared the migration done.

Read the case study →
Open source · Training

MiniAgent

A 104M-parameter advertising model trained from blank weights on a single GPU in 28 minutes for $0.13 of cloud compute, bundled with 14 MCP connectors and a skills framework. Apache 2.0.

Read about MiniAgent →
Article

Claude for marketing teams

John on agents, MCP and the automations that actually hold up in a marketing team's daily work.

Read the article →
Tutorial · Video

Building your first AI agent

From a single script to an agent that calls tools, with the guardrails added at each step.

Read the tutorial →

Frequently asked

What is an MCP server, and do I need one?
MCP, the Model Context Protocol, is an open standard for connecting AI applications to external systems. An MCP server exposes your system's actions as tools any compatible client — Claude, ChatGPT, VS Code, Cursor — can use. You need one if you want AI tools to work with your data without building a separate integration for each.
Which AI model do you use?
Whichever performs best on your eval set for the cost and latency you can accept. We build model-agnostic; our open-source agent supports Claude, GPT and Gemini.
How do you stop an agent from doing something harmful?
Narrow tools, least-privilege credentials handled in code, dry-run previews, human approval on high-risk actions, audit logs, spend caps and adversarial testing — following OWASP's guidance on prompt injection and excessive agency.
Can prompt injection be fully prevented?
Not reliably, and OWASP says so. That's why we design so a successful injection can't do much: the agent reading untrusted content has no write access, irreversible actions wait for a person, and every tool call is logged.
Can an agent work with software that has no API?
Yes, through a browser bridge that drives the admin interface and screenshots every step. We used this to edit CRM workflows in a builder with no public API during the Squeaky Clean Turf migration.
How do you know if the agent is working?
An eval set of real tasks with expected outcomes, run on every change, plus production logs of every tool call and the business metric we agreed at the start.
Will you open-source the agent you build for us?
Only if you want to. Client work is your IP by default. We do bring our own open-source components into builds to save time.

Tell us about the workflow

Describe the task you'd like an agent to take on, the systems it touches, and how you'd measure success.

John, Kristy, or Sandeep will reply. One of the three of us will respond personally within 1 business day. No SDR queue.
We respond within 1 business day. No spam, ever. Read our privacy notice.

Top 25 references

The primary sources, standards, research, and tools we rely on for this work. Every link was checked on 2026-10-11. We aren't affiliated with these publishers unless noted.

Official documentation

  1. What is the Model Context Protocol (MCP)? — Model Context Protocol
    The project's own overview, listing support in Claude, ChatGPT, VS Code, Cursor and other clients.
  2. Security Best Practices — Model Context Protocol
    Known MCP attack patterns, such as confused-deputy and token passthrough, with the mitigations the spec expects.
  3. Tool use with Claude — Claude Platform Docs (Anthropic)
    How Claude decides to call tools, how tool schemas are defined and how results are returned.
  4. Function calling — OpenAI API Docs
    OpenAI's tool-calling interface, including strict schemas that keep arguments valid.
  5. Introduction to function calling — Google Cloud Documentation
    Gemini's function-calling model on Google Cloud, useful when comparing tool reliability across vendors.
  6. Cloudflare Agents — Cloudflare Docs
    The Agents SDK for stateful agents on Workers and Durable Objects, the runtime we usually deploy to.
  7. Build a remote MCP server — Cloudflare Docs
    Step-by-step deployment of an authenticated MCP server that any remote MCP client can reach.

Standards & policy

  1. MCP Specification (2026-07-28) — Model Context Protocol
    The current protocol specification for tools, resources, prompts and transports.
  2. MCP Authorization — Model Context Protocol
    How remote MCP servers use OAuth so agents never handle a user's raw credentials.
  3. LLM01:2025 Prompt Injection — OWASP Gen AI Security Project
    Defines direct and indirect prompt injection and lists the layered mitigations we design to.
  4. LLM06:2025 Excessive Agency — OWASP Gen AI Security Project
    Names the three root causes (excessive functionality, permissions and autonomy) that narrow tools are meant to prevent.
  5. OWASP Top 10 for LLM Applications — OWASP Gen AI Security Project
    The full list of LLM application risks, from sensitive data disclosure to unbounded consumption.
  6. AI Risk Management Framework — NIST
    The U.S. government framework for governing, mapping, measuring and managing AI risk in an organization.
  7. NIST AI 600-1: Generative AI Profile — NIST
    Applies the AI RMF to generative AI, with specific actions for information security and human oversight.
  8. MITRE ATLAS — MITRE
    A knowledge base of real adversary tactics against AI systems, useful for writing adversarial eval cases.

Research & studies

  1. ReAct: Synergizing Reasoning and Acting in Language Models — arXiv (Yao et al.)
    The paper behind the reason-act-observe loop most tool-using agents still follow.
  2. Not what you've signed up for: Compromising real-world LLM-integrated applications with indirect prompt injection — arXiv (Greshake et al.)
    The study that showed instructions hidden in web pages and documents can take over an agent.
  3. τ-bench: A benchmark for tool-agent-user interaction in real-world domains — arXiv (Yao et al.)
    Measures whether agents follow business rules reliably across repeated trials, the reason we test pass rates, not demos.
  4. AgentDojo: A dynamic environment to evaluate prompt injection attacks and defenses for LLM agents — arXiv (Debenedetti et al.)
    A benchmark of realistic agent tasks with injection attacks, useful as a template for adversarial evals.
  5. Defeating prompt injections by design (CaMeL) — arXiv (Debenedetti et al., Google DeepMind)
    Separates trusted control flow from untrusted data, the same principle as keeping write credentials away from untrusted content.

Expert guides

  1. Building effective agents — Anthropic Engineering
    Anthropic's argument for simple, composable workflows over complex frameworks, and when a true agent is warranted.
  2. Writing effective tools for agents — Anthropic Engineering
    Practical advice on tool naming, scope, response size and evaluation that matches how we design tool catalogs.
  3. A practical guide to building agents — OpenAI
    OpenAI's guide to picking agent use cases, designing tools and layering guardrails with human intervention.
  4. The lethal trifecta for AI agents — Simon Willison
    Why an agent with private data, untrusted content and an outbound channel at once is exploitable.

Communities & courses

  1. Model Context Protocol servers — GitHub (modelcontextprotocol)
    Reference MCP server implementations and a directory of community servers to learn from before building your own.
AI disclosure: This page was drafted with AI assistance and edited by a human. Third-party facts link to the official source they came from (checked 2026-10-11); platform names and logos belong to their owners and do not imply a partnership or endorsement.