AgenTomte

September 4, 2026 · 6 min read

AI Agency Partnership Due Diligence: Ask for Evidence, Not Decks

By Anna, co-founder, build and content

Due diligence on an AI development partner is an evidence exercise, not a questionnaire. Before a software reseller signs, four things need an answer you can inspect: whether the partner runs an agent in production today, what happens to AI-written code before it reaches your client, who carries the regulatory role once the system ships, and which parts of the deliverable can actually be assigned onward. Anything answered with a deck instead of an artifact is not answered.

Most of our own work is direct with the company that owns the problem. Partner and white label delivery runs alongside that, and the checks below are the ones we expect to have run on us.

The market has a naming problem, so verify the noun first

Gartner’s press release of 25 June 2025 gave the practice a name: “agent washing”, the rebranding of existing products such as AI assistants, robotic process automation and chatbots without substantial agentic capability. In that release Gartner put the number of genuine agentic AI vendors at roughly 130, out of thousands using the label. It returned to the subject on 20 May 2026, warning that agent washing relabels conventional automation as agentic and raises the risk of misaligned investment and long-term lock-in.

For a reseller that is a supplier classification problem before it is a technology problem. The word on a partner’s website carries no information. A system that has run unattended, on a schedule, against real data, for longer than a sales cycle carries quite a lot.

So ask for one by name. Not a screen recording: the job it does, the date it went live, how often it runs, what it is not allowed to do, and what happens on the runs that fail. A partner who has one answers that in a paragraph. A partner who does not offers a call.

Ask what happens to AI-written code before it reaches your client

Every AI studio writes code with AI now. The difference between them is what runs after the model finishes. Veracode’s Spring 2026 GenAI Code Security Update, published 24 March 2026, tested more than 150 large language models across 80 coding tasks, four languages, and four vulnerability classes, using its own static analysis to grade the output.

Two findings matter to a reseller. Syntax correctness now exceeds 95%, and only 55% of generation tasks produced secure code. Java passed 29% of the time. Cross-site scripting passed 15% and log injection 13%, while SQL injection passed 82%.

Read those together and the commercial risk is clear. Generated code almost always looks finished, and is secure a little over half the time. The gap between “looks finished” and “is safe” is the exact gap your client will never see and your agency will be asked to answer for.

The question is therefore narrow: what runs against generated code before it merges, and who reads the output. A partner with a real answer names the scanning step in their pipeline, the human who reviews, and the languages they avoid. A partner without one talks about the model they use.

Ask who is the provider and who is the deployer

If either party touches the EU, the roles under the AI Act need to be assigned in writing before delivery, not after. The European Commission confirmed on 27 July 2026 that the AI Omnibus had entered into force, moving the rules for Annex III high-risk systems to 2 December 2027 and for high-risk AI embedded in physical products to 2 August 2028.

That deferral is time to allocate responsibility, not a reason to skip the question. It matters more in a reseller arrangement than a direct one, because the client sees your agency’s name on the system and may reasonably treat your agency as the party that put it on the market.

Ask the partner to state, dated and in writing, which AI Act category they believe the build falls into and why, and which of you carries each obligation. “It does not apply” is an acceptable answer only when it comes with the reasoning attached.

Ask which parts of the deliverable can be assigned at all

Ownership has two separate questions in it, and most partnership checklists only ask the first. Whether rights were validly transferred to you is a contract question, covered in what makes the handover safe. Whether the rights exist in the first place is not.

The US Copyright Office released Part 2 of its report, on copyrightability, on 29 January 2025. Its conclusion was that prompts alone do not give a human enough control for the resulting output to be humanly authored, and that protection attaches instead to human-authored material perceptible in the output, and to creative selection, coordination, arrangement, or modification of it.

For code that is usually fine, because the human contribution in a real build is substantial and specific. For material generated wholesale as part of a build, generated copy, images, synthetic datasets, it is a live question. Ask the partner to mark, per deliverable in the statement of work, what they warrant as assignable and what they do not. Drawing that line is a sign they have read the report. Refusing to draw it is also information.

What you askThe answer that failsThe answer you can verify
Show me one agent in productionA demo environment or a recorded walkthroughA named job, its live date, its schedule, its run log
How is generated code checkedThe name of a model or an IDEA scanning step in the pipeline and a named human reviewer
Who holds the AI Act obligations”It does not apply to us”A written, dated category assessment and a split of duties
What can we assign to the client”Everything, it is work for hire”A per-deliverable list of what is warranted assignable
What does it cost if scope moves”We will work with you on that”A published price and a change treated as a new order

What an agency gets when it runs these checks on us

Our own group’s fleet stood at 41 registered agents with 9 live in production as of 30 June 2026, and over the 13 days measured to 4 July 2026 those agents completed 5,450 runs with zero failures. Each one is a tomte, our word for a single production agent with one defined job.

Inspect the discipline rather than the count. Every agent was registered before it ran, has a named owner and a defined fence, writes a run log for every action, and cannot send anything outward without a human approving it. The measurement, including what “zero failures” does and does not mean, is on the agent fleet proof page.

Our prices are published, so a reseller can quote a margin before speaking to us. The AI Workforce Sprint is from $9,500 fixed over four to six weeks, quoted precisely from a written scope, built on the client’s own accounts and handed over with documentation and 30 days of fixes.

Send the scope you have, however rough. You get a fixed number and a delivery window in writing within one business day, and no call is needed at any point. Start async.

Tell us what you want automated

Describe the work in writing. You get a written reply within one business day: a fixed-price proposal, a scoping question, or an honest referral out.

Start at /start

▸ written reply within one business day · no call scheduled, ever

Doesn't fit a package? Tell us what you need anyway.

Questions? Ask in writing

no chatbot · a human replies

Ask us anything, in writing

A founder replies within one business day. That is the same promise clients get.