Here’s a pattern I think a lot of teams will recognize. Take Tarnwell Cycles, the made-up bike brand we use for examples. Its AI visibility report shows that ChatGPT names Aldren first on 6 of the 8 questions about commuter e-bike range. Tarnwell is named on the other 2, after a note about weak cold-weather range. The report even says what to do: publish tested cold-weather range figures on each model page.
A month later, the model pages haven’t changed. Nobody disagreed with the recommendation. The content lead saw it, agreed and filed a ticket. The ticket sat in a queue behind a site migration. By the time a writer picked it up, the context was gone: which questions, which engine, what the answers actually said. So the writer did a general blog post about riding in winter, and the next report looked the same.
Nobody on the team did anything wrong. There were just too many handoffs between the report and the page. This article is about cutting those down by giving the data to the agents your team already uses, while a person still approves every change.
The handoff is where it stalls
Count the steps between “we lose this question” and “we win it”:
- Someone reads the report and spots the gap.
- Someone decides what to change, and on which page.
- Someone writes the change.
- Someone approves it and publishes it.
- Someone checks the next report to see if it worked.
At a lot of companies that’s three people or more across two teams, and each step waits for a meeting or a ticket. None of the steps is hard. Most of the time goes into waiting between them.
Agents are good at the middle of that list: reading a lot of text, pulling out a pattern and drafting something. They’re not good at deciding what your brand should claim, and they shouldn’t publish anything on their own. So the goal isn’t to automate the whole thing. It’s to make steps 1 to 3 take an hour instead of a month, and keep a person on step 4.
Agents can now get the data and steps
Not long ago, getting report data into an agent meant exporting a file and pasting it into a chat. Two things changed that.
MCP gives the agent the data. The Model Context Protocol is a standard way for an AI app to call tools on a server. You point your client at a server URL with a key, and the agent can ask for what it needs. It spread quickly. When Anthropic donated MCP to the Linux Foundation’s new Agentic AI Foundation in December 2025, it counted more than 10,000 active public MCP servers and listed ChatGPT, Cursor, Gemini, Microsoft Copilot and VS Code among the products that had adopted it (Anthropic, December 2025 (opens in a new tab)).
Skills give it the steps. An agent skill is a folder with a SKILL.md file in it: a name, a description and instructions for one task, sometimes with scripts or templates alongside. Anthropic created the format and released it as an open standard, and many agent products now support it (agentskills.io (opens in a new tab)). The agent only loads a skill when a task matches it.
You need both. With only the data, the agent reads the report back to you in nicer sentences. With only the instructions, it gives advice that could apply to any brand. With both, it can pull your lost questions, read what the engines actually said and draft a fix in your buyers’ own words.
The loop worth automating
This is the loop I’d set up first. It’s small on purpose.
- ReadThe latest run for one monitor.
- Find the lossesQuestions where a competitor is named and you aren’t.
- DraftA brief or a page change in the buyer’s words.
- ApproveA named person checks facts and tone, then publishes.A person
- Measure againSame questions after the change ships.
Some things stay human however good the agent gets: publishing, factual claims about your product, pricing, legal wording and the tone of the page. I’d also leave the choice of which gaps matter with a person. An agent will happily draft a fix for every lost question, including the ones you lose because you don’t sell that product.
For Tarnwell, a good draft from step 3 would look something like this: a short “Range in cold weather” section for each model page, with the tested figure and the conditions behind it, written to answer the question buyers actually asked. Under it, for the reviewer, the agent lists the questions the change is meant to fix, the answers that named Aldren instead and the pages those answers cited. The reviewer can check the claim against the range data and see why the change matters, without opening the report.
Three ways to run it
You don’t need to start with a scheduled agent. Here are three setups, from the least work to the most.
-
A chat in your MCP client
Connect Claude Code, Cursor or VS Code to your report data and ask in plain words: “Where did we lose to Aldren this week, and what did the engines say?” You see every tool call, and nothing happens unless you ask.
-
Skills on a weekly routine
Chain a weekly brief, a competitor gap analysis and content briefs for the questions you lost. Run them yourself on Monday morning, or from your own scheduler. What comes out is a document for a person to read.
-
A scheduled agent that drafts
In Contentstack Agent OS, your scheduler calls a trigger and the agent reads the report over MCP, creates draft entries in your stack and posts a Slack digest of what it drafted and what it skipped. It never publishes.
If you’re not on Contentstack, the first two work with any CMS. The agent writes a brief or a draft and a person puts it into your CMS the usual way, or the agent uses your CMS’s own MCP server if it has one.
Guardrails before you hand over a key
Letting an agent read your visibility data is low risk. Letting it write to your CMS, or spend your monthly prompts, needs more thought.
Treat report text as untrusted. AI answers and the pages they cite are written by other people. If a cited page has text aimed at AI agents, your agent reads it too. OWASP ranks prompt injection first in its 2025 list of risks for LLM apps. It describes an indirect kind, where the instructions arrive in outside content such as websites, and it recommends least privilege and human approval for privileged actions (OWASP (opens in a new tab)). Tell your agent in its instructions never to follow instructions it finds in answers or cited pages.
Drafts, not publishing. Don’t give the agent a publish tool at all. If the tool isn’t there, nothing can talk the agent into using it.
Keep the instructions where you can review them. An agent’s instructions decide how it behaves, so treat them like code. Keep them in version control or in the agent builder’s history, and change them on purpose, not halfway through a chat.
One key per agent. Name the key after the agent and keep it in an environment variable, not in a chat or a committed config file. Revoke it when you retire the agent. When something odd shows up in the logs, you’ll know which agent did it.
Confirm before anything that spends prompts. Reading a report doesn’t spend prompts. Starting a new run does. Leave the run tool off, or make a person confirm each run.
Don’t re-run to chase noise. AI answers vary from run to run, so one run is one sample. If an agent keeps re-running until the number looks good, you’re looking at noise, not progress. Compare at least two runs before you call a result. There’s more on this in An AI visibility score is a sample.
Where agent projects go wrong
The data points two ways here. The MCP numbers make agents sound like they’re already everywhere. Meanwhile Gartner predicts that over 40% of agentic AI projects will be canceled by the end of 2027 (Gartner, June 2025 (opens in a new tab)).
I don’t think those contradict each other. Lots of teams can use MCP while many of the projects built on it still stall. My guess is that the ones that stall try to automate a whole job at once, something like “an agent that runs our AI visibility program.” The ones more likely to last do one narrow thing that a specific person checks every week.
So start with one monitor and one question, like “where did we lose to our top competitor this week?” Get that working well enough that someone would notice if it stopped. Then add the next step.
Decide up front how you’ll know it worked. For Tarnwell it could be as simple as the number of those 8 range questions where ChatGPT names Tarnwell, checked over the two runs after the new section ships. If that number doesn’t move, the agent still saved time on drafting, but the draft didn’t fix the answer. That’s worth knowing before you automate the next step.
Before an agent gets your AI visibility data
- Pick one monitor and one question. For example, “where did we lose to our top competitor this week?”
- Make a dedicated key. Name it after the agent, keep it in an environment variable and revoke it when you retire the agent.
- Start with read tools only. Leave the run tool off until a person confirms each run.
- Tell the agent what not to trust. It should never follow instructions found in answers or cited pages.
- Drafts only. No publish tool, and a named person who approves.
- Define done. For example, one draft per lost question, with the buyer’s wording and the evidence behind it.
- Measure again after changes ship. Compare at least two runs before you call a result.
- Log every tool call. Read the first ten drafts yourself before you trust the rest.
Where to start
We built Contentstack Canoe with this loop in mind. Its MCP server gives your client your scores, competitors, every engine’s full answers and the sources they cited. It works in Claude Code, Cursor, VS Code and other MCP clients today, and ChatGPT support is coming soon because it needs OAuth sign-in. There are also eight agent skills that follow the loop above, from a weekly brief to content briefs for lost questions. The ones that change content or spend prompts ask first.
If you run Contentstack, there’s an Agent OS blueprint for the agent our own team uses to draft fixes and post a Slack digest. Import the JSON in Agent OS (Agents, then Create Agent, then Import) and you get a draft agent with the HTTP trigger and six tools. You still set up each tool and the trigger with your own key, stack and channel before you publish the agent. It only ever creates drafts. MCP and skills are on Growth and Enterprise.
Whichever way you start, pick the one question you’d most like answered every Monday and build that first.
