AI strategy · · 7 min

How to evaluate an AI agent development company: a checklist

Twelve questions to ask any AI agent vendor, what a good answer sounds like, and the red flags. It ranks no one, and we answer the same questions ourselves.

Choosing an AI agent development company is hard because every vendor’s demo looks good. Demos run on clean data, happy paths and a presenter who knows what to type. Production agents run on messy data, edge cases and users who don’t. This checklist is built to tell those two apart.

It is vendor-neutral on purpose. It ranks no one, it recommends no list, and it works for any team you are talking to. Disclosure: Apptycoons builds AI agents, so at the end we answer the same twelve questions about ourselves, and you can hold us to them.

Why demos are a poor way to choose an AI agent vendor

A demo shows that a model can do a task once; a production agent has to do it thousands of times, safely, with a record of what it did. The questions that separate vendors are about evaluation, permissions, failure handling and ownership, and none of them show up in a demo.

It also helps to know what a good vendor will often recommend: less agent than you expected. Anthropic’s engineering guidance on building agents advises finding the simplest solution possible and adding complexity only when needed, and distinguishes fixed workflows from agents that direct their own steps. A vendor who proposes a simple workflow where it fits is usually a better sign than one who proposes a multi-agent system for everything.

1. Can they show an agent running in production?

The strongest evidence is a live product you can use, or a customer you can call. Ask what the agent does, how long it has been live, what broke after launch and how they found out.

Good answers are specific: a monitoring dashboard, a class of failure they didn’t expect, the fix and the test they added. Weak answers stay at the level of frameworks and model names.

2. Do they start from your workflow, not their framework?

A good vendor asks about the job before the stack: what triggers it, who does it now, what systems it touches, what a correct result looks like and what a mistake costs. If the first meeting is about frameworks and vector databases, the scoping hasn’t started.

Ask them to describe the first agent in one sentence, with one metric. If they can’t, the project will drift.

3. Can they say what the agent will not do?

Scope is defined as much by exclusions as by features. Ask the vendor to write down which actions the agent may take, which it may only propose, and which it must never touch.

This maps to a real security risk. The OWASP Top 10 for LLM Applications 2025 lists excessive agency, an AI system given more functionality, permissions or autonomy than it needs, as one of the ten main risks. The fix is design, not a disclaimer: least-privilege tool access, separate credentials, and approval steps for anything irreversible.

4. How will they evaluate the agent before and after launch?

An agent without an evaluation set cannot be improved safely. Ask how the vendor will build a test set from your real cases, what score counts as good enough to launch, and how every change to a prompt, model or tool is re-scored before it ships.

Good answers include real examples from your data, a pass mark agreed before the build, and regression testing when models change. Any accuracy promise made before seeing your data is a guess.

5. How do they defend against prompt injection and data leaks?

Ask directly how they handle prompt injection, the first risk on the OWASP list, and sensitive information disclosure, the second. Any agent that reads email, documents, web pages or user input can receive instructions hidden in that content.

Look for layered answers: untrusted content treated as data, tools scoped to the minimum, outputs validated before they trigger actions, sensitive data redacted, and model endpoints configured not to retain your data.

6. Where does a human approve, and is every action logged?

Every consequential action should have a clear owner. Ask where in the flow a person reviews or approves, how the agent escalates when it is unsure, and whether every action is logged with its inputs and outputs in a way an auditor could follow.

If your industry is regulated, ask how the logs map to your existing controls. Frameworks such as NIST’s AI Risk Management Framework, a voluntary framework for managing AI risk, give you a shared vocabulary for that conversation.

7. Which models do they use, and can you switch?

Models change every few months, so lock-in to one provider is a cost. Ask whether the architecture is model-agnostic, how a model swap would be tested, and what happens to cost and latency if you move.

A good answer is often “we pick per task”. Combined with an evaluation set, that lets you change models with evidence rather than hope.

8. How will it connect to your systems?

Most of the work in an agent project is integration, not prompting. Ask which of your systems the agent reads from and writes to, how authentication works, what happens when an API is down, and whether the agent can do anything a normal user with the same role could not.

9. Who owns the code, prompts, evaluations and data?

You should own everything that makes the agent work: code, prompts, evaluation sets, infrastructure configuration and documentation. Ask for it in the contract, and ask whether the agent can run in your own cloud account.

Then ask the uncomfortable question: what happens to the agent if the vendor disappears? If the answer involves their proprietary platform, you are renting, not buying.

10. What is fixed in the price?

Compare what is fixed, not the headline number. Ask whether there is a paid discovery phase, whether the build price is fixed after it, how payments map to delivered features, and what running costs (model usage, hosting, monitoring) you should expect.

Open-ended time-and-materials with no discovery is not wrong, but it moves all of the risk to you. Our AI agent development cost guide breaks down what a build and its running costs usually include, so you can check a quote line by line.

11. What happens after launch?

Agents need care after launch: models update, data drifts, and users find new edge cases. Ask what support is included, how monitoring works, who watches cost and success rate, and how improvements are scoped.

12. Will they tell you when not to build an agent?

The best signal of an honest vendor is a recommendation against their own interest. Ask: “What would you not automate here?” A good partner will point to steps where a rule, a form or a person is cheaper and safer than a model, or to an off-the-shelf agent that already does the job. Our build vs buy framework for AI agents is a useful way to test that advice.

Red flags when choosing an AI agent developer

These are the patterns that should end a conversation or at least slow it down:

  • Accuracy or savings figures promised before anyone has seen your data.
  • No evaluation set, or “we test it manually”.
  • No answer on prompt injection or excessive agency.
  • Write access to production systems with no approval step.
  • Prompts, evals or orchestration kept as the vendor’s intellectual property.
  • A demo built on data that is not yours, presented as proof.
  • No discovery phase, and no clear statement of what is fixed.

A scoring sheet you can copy

Score each vendor from 0 to 2 on each line (0 = no answer, 1 = general answer, 2 = specific evidence). Compare totals only after the calls, and weigh the first four lines double.

Vendor scoring sheet (0 to 2 per line)
CriterionWhat a 2 looks likeWeight
Production evidenceLive product or callable referenceDouble
EvaluationTest set from your data, agreed pass mark, regression testingDouble
Limits and approvalsWritten action boundaries, human approval, audit logDouble
SecurityPrompt injection, data leakage and data retention answered specificallyDouble
Workflow fitFirst agent in one sentence with one metricSingle
IntegrationSystems, auth and failure handling explainedSingle
Model flexibilityModel-agnostic, swap tested with evalsSingle
OwnershipCode, prompts and evals yours, deployable in your cloudSingle
Pricing clarityDiscovery, fixed scope, milestone payments, run costsSingle
After launchMonitoring, support and improvement processSingle

Our answers to the same questions

Here is how Apptycoons answers the checklist. Hold us to it, and ask us for the evidence on a call.

  • Production evidence: Relm, a multi-agent system that researches commercial real estate deals with every finding cited to a source record, and Superdeal, an AI agent that runs creator campaigns with the brand approving every step. Both are live products.
  • Workflow first: our process starts by mapping one workflow and one metric.
  • Limits and approvals: approval steps, escalation rules and full audit logs for every action. On Superdeal, every consequential step goes Draft, Review, Send.
  • Evaluation: a test set built from your real cases, scored before launch and again after every prompt, model or tool change.
  • Security: least-privilege permissions, PII redaction and zero-retention model endpoints, so your data does not train public models.
  • Models: we choose per task from Anthropic, OpenAI, Google or open-source models and keep the model swappable.
  • Ownership: the code lives in your repository from day one, the IP is yours, and the agent runs in your cloud or on endpoints you control.
  • Pricing: published on our pricing page. Discovery is two weeks at $5,000, credited to the build; builds are fixed at $25k to $120k; a Dedicated Team starts at $12k a month. Every build includes 30 days of post-launch support.
  • When not to build: if a rule or a form solves the problem, we will say so. Our post on taking AI pilots to production covers why so many stall.

If that matches what you need, our AI agent development page has the details, and a free AI audit is the fastest way to test our answers against your workflow.

Related service

We design and build autonomous AI agents and multi-agent systems that read, decide and act across your tools, with human approval where it matters.

Want this applied to your business?
Free 45-min AI audit with a senior architect.
Book the audit
FAQ

Common questions.

What should I ask an AI agent development company before hiring?

Ask to see an agent they run in production, how they will evaluate yours before and after launch, what the agent will not be allowed to do, where humans approve actions, how they defend against prompt injection, who owns the code and prompts, and what is fixed in the price.

How do I know if an AI agent vendor has real production experience?

Ask for a live product or a reference customer you can contact, then ask what broke after launch and how they found out. Teams with production experience answer with specifics about monitoring, failures and fixes; teams without it talk about frameworks and demos.

What are red flags when choosing an AI agent developer?

Accuracy promises without an evaluation set, no answer on prompt injection, agents with write access and no approval step, prompts or evals the vendor keeps as their property, pricing that is open-ended with no discovery phase, and a demo built on data that is not yours.

Who should own the code and prompts of a custom AI agent?

You should. Ask for the code, prompts, evaluation sets, infrastructure configuration and documentation to be yours in the contract, deployable in your own cloud account, so the agent survives a change of vendor.

How much does it cost to hire an AI agent development company?

Compare what is fixed rather than the headline: a paid discovery, a fixed build price after it, milestone payments and estimated running costs. For reference, Apptycoons publishes a $5,000 two-week discovery, credited to the build, then fixed-price builds of $25k to $120k.