Skip to contentSkip to contact
Ilja Nevolin

Agent factories · SaaS exits · Robotics · Worldwide

Stop renting software your engineers could own.

Agent factories, robot-fleet software, and your own replacements for the SAP and Salesforce modules you overpay for. Built in your repo, run by your team.

Not website chatbots. Not foundation models. Systems that run in production.

36% of SaaS licences go unused, on average.

Zylo, 2026 SaaS Management Index, Jan 2026 (opens in a new tab)

Median SaaS spend is $9,455 per employee. At renewal, the idle third is your budget to own the part you use.

Six systems I build.

01, SaaS exit

Rebuild the part of Salesforce you actually use.

Agents map what you really use, rebuild that slice as your own service, and shadow-run it against the vendor until every response matches.

  • Salesforce pipeline rebuilt on Postgres with an agent service; Service Cloud kept
  • SAP ECC Z-reports and RFC calls moved behind a typed API, so the S/4 decision gets smaller
  • ServiceNow ticket routing replaced by a queue, a classifier and rules your team owns

Hard part: data migration and reconciliation, not the UI. Switch-back kept until parity.

35% of teams have already replaced at least one SaaS tool with a custom build.

Retool, 2026 Build vs Buy Report (817 builders, late 2025) (opens in a new tab)

The CRM module is broken into the objects actually used; agents rebuild deals as a tested service, shadow-run it against the vendor until every response matches, and the first 400 seats are released.

02, Agent factory

One agent is a demo. Twenty need a factory.

Templates, a tool registry, sandboxes, evals and token budgets, so your teams ship agents the way they ship services.

  • Migration agents that move every service to the next framework version, one CI-green PR each
  • Flaky-test triager, dependency bumper and on-call summariser from one template and one MCP registry
  • Every run traced: prompt, tool calls, tokens and eval score, in the observability you already run

Hard part: evals that check outcomes, like tests passing, not wording.

Over a thousand pull requests merged each week at Stripe are written entirely by its coding agents, then reviewed by engineers.

Stripe, "Minions", Feb 2026 (opens in a new tab)

Agents move along four lanes (spec, build, eval, deploy); one fails a prompt-injection eval and overruns its token budget, and the platform lead sends it back with the failing cases.

03, Context layer

Wire each system once, not once per agent.

One MCP gateway over your code, logs and the SaaS you still run: cited answers, scoped tools, production writes behind approval.

  • "Why did crm-sync start failing on Tuesday?" answered from git, API logs and the job queue
  • MCP servers for Postgres, GitHub, Jira and Datadog behind one gateway, scoped per agent
  • Retrieval over ADRs, runbooks and postmortems, citing the commit or page behind each answer

Hard part: permissions. Every tool call runs as the person asking.

An engineer's question is answered from the code, the CRM API log and the ERP job queue with three cited sources; production stays locked; the fix is opened as a pull request that needs approval.

04, App platform

Prompt the screen. Skip the seat.

A prompt-to-app builder on your data, SSO and components: teams describe the tool, your platform reviews and ships it.

  • A pipeline board on the new deals service, replacing the CRM screens reps actually open
  • Prompt to PR to preview environment, on your component library and monorepo
  • Row-level data scopes, so a generated app can't read payroll

Every app lands as a reviewed PR in your repo, with SSO, a data scope and an owner.

A RevOps engineer describes a pipeline board; it assembles on the deals service and ERP orders, lands as a reviewed pull request with sign-on, a scoped role and an owner, and replaces the CRM screens reps used.

05, Robotics and fleets

Twelve robots slow down. One cause, found overnight.

Fleet telemetry and bag replays fold into one incident, with a reversible fix staged before on-call wakes.

  • An LLM task layer over ROS 2 through MCP, inside a safety envelope the model cannot override
  • Teleop data in LeRobot format and sim evals in Isaac or MuJoCo, gating every policy before hardware
  • Line and PLC state read over OPC UA: agents read, never write setpoints

The model never sits in the safety loop: agents may slow, reroute or close a zone from your playbook; firmware waits for a person. I build the agent, data and fleet software; your robotics engineers own the safety case.

A fleet of 96 robots reports overnight; 38 alerts fold into one cause, twelve robots on a firmware canary; an agent slows them and closes one aisle from a playbook, reproduces the drift in simulation and stages a rollback, which the robotics on-call approves at 07:40.

06, AI-first engineering

Writing code faster isn't shipping faster.

You already pay for coding agents. I make them show up in lead time: specs first, eval gates in CI, review agents on every PR.

  • Background agents that land migrations and dependency bumps as reviewable PRs
  • Review agents clear the nits, so senior engineers review the design
  • Robot planner changes gated by simulation scenarios before they reach hardware

Hard part: measuring it. PR count goes up either way.

Teams with high AI adoption: 98% more PRs merged, review time up 91%, no measurable company-level gain.

Faros AI, AI Productivity Paradox Report, Jul 2025 (opens in a new tab)

The crm-sync fix from the context layer fails three named evals, goes back to the agent, passes all 215, is approved by a reviewer and merges; the practice rolls out team by team.

Month 6, What you own

The factory stays. The code is yours.

Agents, tools, evals and apps in your repo and cloud. Your engineers run it and build the next one.

  • Your cloud, repo and CI
  • Your SSO and permissions
  • Every agent action traced
  • A person approves prod, money and fleet changes

An example month 6 on one card: 400 of 1,200 CRM seats released, 14 agents in production, and the CRM licence down from €1.8M to €1.2M a year against €6k a month to run.

How we start

A working spike on your data in two weeks.

Then production, then hand-over. Stop after any step and keep everything built. No licence to me, ever.

  1. 01

    Feasibility sprint

    2 weeks

    What you hold at the end: A working spike on your data, a monthly run-cost estimate, and a go or no-go.

  2. 02

    Build to production

    6–10 weeks

    What you hold at the end: One system live in your cloud: evals in CI, traces, a token budget and a rollback path.

  3. 03

    Hand over

    4–8 weeks

    What you hold at the end: Your engineers own it and build the next one on the same factory. Code, runbooks, evals. No licence to me.

  • Renewal coming up? Start here.

    SaaS exit plan

    Which SAP, Salesforce or ServiceNow modules to rebuild or keep, costed per seat against run cost, with a slice-by-slice cutover plan.

    Ask about an exit plan
  • Already running agents? Start here.

    Agent security review

    Prompt injection, tool misuse and data leaks tested against the OWASP Top 10 for agentic applications, with a fix list ranked by risk.

    Ask about a review
  • Fractional architect

    A senior hand on your agent platform, robotics software or SaaS exit, without a hire.

A work order is stamped SPIKE after two weeks, LIVE once the system runs in production and YOURS at hand-over, when the key passes to your team.

About

You work directly with the engineer who builds it.

Ilja Nevolin, working with teams worldwide. Fifteen years of engineering, from HPC simulation to crypto infrastructure; now building agent platforms at enterprise scale.

  • Now:

    Solutions Architect, agentic AI platforms at enterprise scale; before that, Staff Engineer putting coding agents into its engineering teams.

    What holds up at enterprise scale, this year.

  • Security:

    Pentesting and code review · software-protection research, Ghent University (IEEE EuroS&P workshop, 2020).

    I test for: prompt injection, tool misuse, privilege abuse, data leaks

    Agents your security team can sign off.

  • Built at:

    Notabene (crypto compliance) · bp · Keysight R&D (EDA, FEM, HPC).

    The maths under simulation and planning: FEM and HPC at Keysight R&D.

  • Founded:

    Two startups, acquired 2017 and 2018 · co-founder of Carbon3, climate-finance infrastructure.

    Shipped on startup budgets, twice to acquisition.

LinkedIn (opens in a new tab)
Ilja Nevolin, portrait

Contact

Send me the system you'd rather own.

Three lines is enough. You get a 30-minute call and a straight answer: feasible or not, rough cost to build and run, first step.

Reply within two working days.

Opens your mail app. Nothing is sent from this page.

The contact card fills in as you type: what to build, what it runs on today and your team.