AgenTomte

September 10, 2026 · 7 min read

Our AI Product Development Process, From Brief to Launch

By Sahan, co-founder, systems and delivery

We run new product development in four stages: scan the category, write the brief, generate the technical specification, then hold a human gate before anything moves. An agent drafts at every stage. A person makes every decision that spends money or commits a formulation.

That split is the design. Drafting is where the hours go. Deciding is where the risk sits, and no agent of ours is allowed near it.

The shortlist is worth more than the idea

NielsenIQ published “Launch Fast, Learn Faster” on 20 April 2026, covering more than 3,500 new brands launched in the UK, Germany, France, Italy and Spain during 2025. Roughly one in three reached 1% household penetration. Only 17% achieved both strong trial and strong repeat purchase. Perishable food was the strongest category at 37%. NielsenIQ measured this against its retail purchase panel and interviewed more than 54,000 verified buyers, so it is purchase behaviour rather than opinion.

You will also see a claim that 95% of new products fail. We do not use it. Clayton Christensen, to whom it is usually attributed, denied saying it, and the nearest thing to a primary source, Castellion and Markham in the Journal of Product Innovation Management in 2013, puts real consumer goods failure closer to 40 to 49%. Two thirds of launches missing 1% penetration is a big enough problem without inflating it.

The number that matters for process design is the first one. Most launches are decided before anything is formulated, in the choice of what to develop at all. So that is where our pipeline starts.

Stage one: the scan runs weekly and starts from evidence

Our opportunity-scanning agent sweeps public sources on a weekly schedule: trade press, retailer new-arrival pages, search trend data, regulator feeds. It groups raw signals into themes, writes each theme up as a structured opportunity card, and scores every card on a fixed weighted rubric covering brand fit, evidence strength, demand signal, feasibility, differentiation and regulatory clarity. The weights sit in configuration rather than in code, so they can be retuned once real runs show which ones were wrong.

Two rules constrain it. Every card carries at least one real, verifiable source link or it is discarded rather than shown. And anything already in the catalogue, or already surfaced in recent weeks, is dropped before scoring, so the same idea cannot keep arriving dressed as new.

Stage two: a human clicks before any work starts

Ranked cards land in the brief-writing agent’s intake queue automatically. The delivery is automatic. The click that turns a card into a project belongs to a person, and nothing downstream happens without it.

Once clicked, the agent structures the request into objectives, target market and special requirements, then runs the supporting research: benefits, actives, competitive landscape, technical feasibility, regulatory considerations, and a recommended development pathway with its risks written out.

Stage three: the spec says target until a lab says otherwise

Next it drafts two documents. A manufacturing document with the formulation table, equipment list, step by step process and critical control point tables for things like moisture and foreign metal control. Then a technical specification: organoleptic, chemical, heavy metals, microbiological, mycotoxin, packaging, shelf life and a compliance statement.

Every figure in both documents is labelled as an R&D target pending lab confirmation. The agent has no way to mark a value confirmed. That transition only happens when a human enters an accredited lab result.

The labelling does real work. When we queried the FDA’s openFDA food enforcement database on 10 September 2026, it held 1,617 food recall records with 2025 report dates. Those are product and distribution level records rather than deduplicated recall events, so read it as an indicator of scale, not an event count. Specifications are where that class of cost gets decided, months before anyone ships a pallet.

Who does what, at each stage

StageWhat the agent producesWhat a human must do
ScanRanked opportunity cards, each with a cited sourceNothing. Read the shortlist or ignore it
BriefObjectives, target market, requirements, research reportClick to start it. Approve the direction
SpecManufacturing document and technical specification, all values marked targetEnter lab results. Sign off the formulation
TrackNothingMove the card, set the launch date, call go or no go

Four things that broke

Silent truncation. Our specification generator quietly stopped around section 3 of 11 because the output length allowance was too small. No error appeared. The document simply ended, and it read as finished if you did not scroll to the bottom. We raised the ceiling twice before the whole spec, heavy metals and shelf life included, generated intact. If you build a document pipeline, size the output ceiling for the longest document you will ever ask for, and make sure somebody actually reads the last section.

Fabrication under pressure. Before this scanner existed we ran the same job on a no-code automation platform. Steered interactively through its chat interface, it invented product launches that had never happened and described them convincingly. That failure is exactly why every opportunity card now has to cite a real source or die.

A rename broke the cost model. In our costing tool, cost inputs link to their source materials by name rather than by a stable identifier. Renaming one material silently broke the calculation for everything downstream of it, with no error raised anywhere. The costs simply stopped updating. We have patched the visibility, so unlinked items get reported for a human to catch, but the underlying design gap is still open. We are publishing this one while it is still unfixed.

Deployed is not used. More than one of these tools reached a built and deployed state while the human step that closes the loop had never once run in production. Nobody had clicked generate brief on a queued card. The deployment status said the system was live. The run logs said nobody had touched it. We check the run logs now before anyone calls a thing finished.

Where our process still stops short

This post’s title could imply we generate a launch checklist. We do not. Launch readiness sits on a kanban board as a human-managed status, R&D through sampling, production and commercialised, with an on-hold state that requires a written reason, plus a human-set target launch date. A person moves the card. The system never advances a stage on its own, and we have not tried to make it.

Three more gaps worth naming. Costing lives in a separate agent and does not appear in the generated document set at all. The catalogue check that stops the scanner re-suggesting products we already sell is a hand-maintained list rather than a live feed from the storefront. And the scanner reads public sources only. It cannot yet see our own first-party demand signal: the support questions, the on-site searches that return nothing useful, the reviews. That is the single biggest improvement available to us and it is not built.

Why smaller companies reach this before big ones do

The US National Science Foundation published its Business Enterprise R&D survey results on 29 September 2025. Companies with 10 to 249 employees accounted for 9% of the $722 billion that American businesses spent on R&D, while carrying the highest R&D intensity in the survey at 11% of sales, against 5% for large firms. Smaller companies already spend a larger share of their revenue developing products. They just do it with far fewer people, which is the condition this pipeline was built for.

The AI is not going there yet. The US Census Bureau’s AI supplement, published 26 May 2026, put firms of 100 to 249 employees at 32% AI use, with 57% of adopters running it in three or fewer business functions, concentrated in sales and marketing.

Eurostat’s figures, extracted in December 2025 from a compulsory survey of roughly 157,000 enterprises with 10 or more staff, show R&D and innovation as a leading AI application in one sector only, information and communication, at 42.51% of AI-using enterprises there. Those are self-reported, so read direction rather than precision. Everywhere else, the agents went to marketing first, and product development is still open ground.

If you want the shorter version of how we decide what to build at all, the ranked-list method in our AI audit checklist is the same logic applied to processes instead of products, and product costing automation covers the pricing agent that sits downstream of this one.

What a build like this is priced at

An AI Workforce Sprint is from $9,500 for one working agent or internal application, delivered in four to six weeks, shipped to your own accounts with documentation and a recorded handover. The pipeline described here is two agents plus a dashboard, which is two sprints rather than one.

We run every agent in the group the same way, with the same rule about what a machine is allowed to decide. The agent fleet case study carries the dated numbers: 41 registered agents, 5,450 runs and zero failures across 13 days, as of July 2026.

If new product development in your company currently lives in one person’s head and a spreadsheet, write it down and send it to us. Two minutes at /start, written reply within one business day, no call at any point.

Tell us what you want automated

Describe the work in writing. You get a written reply within one business day: a fixed-price proposal, a scoping question, or an honest referral out.

Start at /start

▸ written reply within one business day · no call scheduled, ever

Doesn't fit a package? Tell us what you need anyway.

Questions? Ask in writing

no chatbot · a human replies

Ask us anything, in writing

A founder replies within one business day. That is the same promise clients get.