Legacy ERP automation with computer use agents: how it works and how to set it up
A computer use agent is an AI model that works a computer the way a person does. It reads a screenshot and decides the next click or keystroke, and a program on a dedicated machine carries it out. That lets it key orders and invoices into a legacy ERP (older order, stock and accounting software installed on your own server) through the same screens your staff use. You need no API (a published way for one program to exchange data with another) and no new ERP. It makes sense for the entries that no import file, database connection or API can make, with a named person approving each entry before it is saved.

What a computer use agent does on an old ERP
Most AI tools reach business software through an API. Many older ERPs have none, or none that covers the screen you need, so a computer use agent works the screen itself. Anthropic offers a computer use tool for Claude, generally available on its own API. OpenAI offers computer use for its GPT models. Google’s Computer Use for Gemini is a preview, and Microsoft offers computer use in Copilot Studio, its tool for building AI agents.
There are also open-weight models, whose files you can download and run on your own server. Holo3-35B-A3B from H Company and UI-TARS-1.5-7B from ByteDance are built to operate on-screen interfaces, and both are published under the Apache 2.0 licence.
a16z, a venture capital firm, wrote in August 2026 that computer use agents are beginning to hold up in production on narrow, repeatable workflows. Its examples include updating systems of record, the systems that hold a company’s official records, and handling software where no clean API exists. An old ERP without an API fits both examples.
How the loop works
Anthropic calls this cycle “the agent loop.” In plain words:
- A control program, which you run on your side, sends the model the task and a screenshot of the ERP screen.
- The model reads the screenshot and replies with the next actions, such as a click on the customer field, the customer number to type, or a press of the Tab key.
- The control program performs those actions on a dedicated Windows machine inside your network.
- It takes a new screenshot and sends it back, so the model can see the result.
- The cycle repeats until the task is done, a step limit is reached, or the control program pauses for a person to approve.
The control program is where your rules live. It decides which actions run and when to stop for a person. OpenAI’s computer use integration guide says it plainly: “The model’s request to act is not user permission.” Anthropic’s computer use documentation says the screenshots and keystrokes of a session are stored in your environment, and that it processes them in real time as part of each request. Where that processing happens is covered in where the screenshots go.
Which old ERPs it works on
An agent that reads screenshots can, in principle, work any screen a person can work. Anthropic warns that reliability can be lower with niche applications, and an old ERP is a niche application. How well it goes depends on how the ERP reaches the screen.
Windows desktop clients
Many older ERPs run as a Windows program installed on each user’s PC. Windows includes Microsoft UI Automation, an accessibility framework built for tools such as screen readers, which gives programs access to most of the buttons and fields on the desktop. Microsoft notes that automated test scripts use it too. Where an ERP screen exposes its fields this way, a script can fill them directly, which is faster than an agent and does not depend on reading a picture.
Some of these products are near the end of support. Microsoft ends Dynamics GP support on December 31, 2029, and will make security updates available, if needed, until April 30, 2031. An agent does not change those dates. It can take retyping off your staff while you decide what replaces the ERP.
AS/400 and green screens
IBM i, the system many people still call the AS/400, is often run through green screens: text-only screens reached through a 5250 terminal emulator (a program on a PC that stands in for the old IBM terminal) and driven from the keyboard. They suit an agent better than they look, because everything on them is text. HLLAPI is a programming interface that lets another program read a terminal screen’s text and send keystrokes. Microsoft’s Power Automate documentation says it is “supported by nearly all terminal emulation software,” and Power Automate for desktop, Microsoft’s tool for automating work on a Windows PC, uses it to get text from a terminal session. After the agent types, the control program can read the exact characters in each field and check them without relying on the screenshot.
ERPs run through Citrix or Remote Desktop
Citrix and Remote Desktop run the ERP on a central server and send a picture of its screen to your PC. An agent works from that picture as a person does. Plan the sign-in early: Microsoft’s Copilot Studio documentation says the passwords it stores for computer use might not work in Citrix or other virtualized environments. Test one remote session before you commit to a design.
The routes into an old ERP, compared
Screen work is the slowest route into an ERP and the hardest to predict, so give an agent only the steps that nothing else can do. Check the other routes first. The reliability and speed columns give our own reading of how each route behaves.
| Route | What it needs | Reliability | Speed | When to use it |
|---|---|---|---|---|
| Read-only database login | A database account that can only read, such as SQL Server’s db_datareader role, or an IBM i ODBC connection (a standard way for programs to query a database) set to read-only | High, with exact values | Seconds | Every lookup: prices, stock, open orders, customer part numbers and history |
| Import files | The ERP’s own import tool, such as Sage 300’s batch import | High, when each batch is checked | Minutes per batch | Entries in volume that the ERP already accepts from a file |
| EDI | Electronic data interchange, a standard document format exchanged with trading partners, and EDI software mapped to your ERP | High once mapped | Minutes | Customers or suppliers who already send orders or invoices by EDI |
| Vendor API | An API for your ERP version, and its licence terms | High | Seconds | When your version has one that covers the screen you need |
| UI automation | A Windows screen that exposes its fields to UI Automation, or a green screen reached through HLLAPI | High while the screen stays the same | Fast | Stable screens with known steps |
| Classic RPA (robotic process automation) | An RPA tool, which replays recorded clicks and keystrokes, plus agent steps where a screen varies | High on the recorded part | Fast, slower on agent steps | Routine flows with occasional exceptions |
| Computer use agent | An isolated machine, a model, a control program and an ERP licence for the agent | The lowest here, so measure it on your screens | Slower than a person | The steps no other route can reach, with a person approving each save |
Read and write through different routes. Read through a read-only database login wherever you can, because it returns exact values in seconds and cannot change anything. Write through the ERP’s own screens or import tools, so that the ERP’s own checks, such as required fields and credit limits, still run. Any write route needs a check afterward. Sage’s import help warns that “incorrect importing can damage your data,” and says to print the batch listing after each import and check it against the original documents.
Computer use agents and RPA on legacy systems
RPA (robotic process automation) is software that replays recorded clicks and keystrokes. It is fast and repeats exactly, and it tends to fail when a screen changes or an input does not match the recording. An agent copes with variation better and costs more per step.
UiPath, an RPA vendor, recommends combining the two. Its ScreenPlay best practices set the goal as “a maximally deterministic, minimally agentic workflow”: recorded steps wherever the screen is stable and the steps are known, and agent steps only where they are the only reliable way to meet the business need. The same page warns that a workflow built on exact screen selectors (the fixed addresses a script uses to find each button and field), which breaks with every minor change, can cost more over its lifetime than the model fees it saved. For an old ERP, that usually means a recorded path for the routine order and an agent for the exceptions, such as an unexpected warning or a field that moved.
Five ERP jobs to automate first
Start with a job your team repeats every day, where each entry can be checked against a source document.
- Order entry from emailed purchase orders. A customer’s PO arrives as a PDF, and someone keys it line by line. The worked example below follows this job. Reading the PO itself is document extraction, covered in text extraction from images.
- Price and stock lookups for quotes. These are reads, so they belong on a read-only database login. The agent needs the screen only if a price is calculated on screen and stored nowhere you can query.
- Supplier invoice keying. The agent matches the invoice to its purchase order and receipt, and enters it for approval. It stops before posting, the step that makes an entry final in the ledger.
- Month-end report pulls. The agent runs the same reports with the same filters each month and saves them to a folder. Nothing is written to the ERP, so it is a good first job to learn on.
- Customer part number lookups. The agent translates a customer’s part number to yours from the ERP’s cross-reference, and flags it when no match exists.
For the other jobs AI can take at each desk in a plant, see AI for manufacturing, module by module.
A worked example: purchase orders from email
The company, volumes and steps below are made up to show how the pieces fit. They do not describe a client project.
Picture a distributor with 60 employees and an older Windows ERP. Customer service keys about 40 emailed purchase orders a day into the order entry screen. With an agent in place, each order goes like this:
- A PO arrives in the orders mailbox as a PDF. A document-reading step pulls the customer, PO number, ship-to address, lines, customer part numbers, quantities and requested dates into a draft.
- The control program checks the draft against the ERP through a read-only database login. It matches each customer part number to your item number, and it looks for an existing order with the same customer PO number.
- A customer service rep sees the draft beside the PDF, fixes anything flagged and approves it.
- The agent opens order entry on its own virtual machine, a computer simulated in software on your server, and types the header and each line. Before it may save, the control program compares every field on screen with the approved draft.
- After the save, the control program reads the new order back through the read-only login and compares it with the approved draft again. Any difference stops the run and sends the order to the rep.
- The log keeps the PDF, the approved draft, every screenshot and every keystroke, filed under the PO number.
The rep’s job changes from retyping each order to reviewing a draft beside the PDF. Whether that saves time depends on how many drafts come back clean, and you only learn that by measuring on your own orders.
Which ERP screens does your team retype?
Tell Derik which ERP you run and which entries your team keys by hand. He will tell you which route fits each step, and where an agent makes sense.
Start a conversationHow to set up a computer use agent on an old ERP
Reads come first, then writes one screen at a time, with a measured pilot before anything touches live data.
- Pick one job and write down the procedure. Choose a high-volume, low-risk entry, such as order entry from emailed purchase orders. Write down each screen, field and known warning the way you would for a new hire.
- Move every read to a read-only database login. Give the control program a database account that can only read. Use it for lookups and for checking each entry after it is saved.
- Build a test company. Copy the live company into a test company, or use the sample data your ERP ships with, and run every early test there.
- Set up an isolated Windows virtual machine. Install only the ERP client, fix the screen size, and let the machine reach only the ERP server, the password vault and the model provider.
- Create the agent’s ERP user and licence. Give the agent a user of its own with rights to the screens its job needs. Where your ERP separates saving from posting, let it save only. Check that the user is licensed.
- Put the password in a vault. The control program fetches the password and signs in before the agent starts, so the password never appears in the model’s instructions.
- Set limits and a known starting state. Start every task from the same screen, set a maximum number of steps and a time limit, and write the known pop-ups and warnings into the procedure.
- Add the approval step. A named person approves each entry before the agent saves it, and the control program refuses any save whose screen does not match what was approved.
- Add the read-back and duplicate checks. Check a unique reference, such as the customer’s PO number, before creating a record. After each save, compare every field with the approved entry.
- Keep a log. Store each screenshot, action and approval under the business reference, with the same access rights as the ERP and a fixed retention period.
- Pilot 50 to 100 transactions in the test company. Measure the share that finish correctly without help, the minutes and model cost per transaction, and how often a person had to step in.
- Choose where the model runs. Decide between a model on your own server and a hosted model under a written zero data retention agreement, under which the provider does not store your requests or its answers after it replies, before any live data reaches it.
Then move to live data one screen at a time, starting with the job the pilot handled best. For how an agent project is scoped and staffed beyond the ERP, see how to build an AI agent for a manufacturer or distributor.
Best practices for an agent on your ERP
Run it on an isolated virtual machine
A virtual machine is a computer simulated in software on a server, which you can reset or replace without touching anything else. Anthropic, OpenAI and Google each recommend running computer use in an isolated virtual machine or container (a lighter, sealed-off software environment). Microsoft’s Copilot Studio guidance asks for dedicated machines used only for computer use, with only the essential applications installed and web access limited to an allow list, a list of the only addresses the machine may reach. For an ERP agent, that means a Windows virtual machine inside your network with the ERP client, a connection to the model and the password vault, and nothing else. Google adds that unexpected pop-ups and notifications confuse the model, and recommends starting each task from a known, clean state.
Start read-only, in a test company
Run the first weeks on reads alone, such as reports, lookups and the checks that compare a document with the ERP. When writes begin, run them in a test company, a copy of your company data kept apart from the live books, or in the sample data some ERPs ship with, as Sage 300 does. Move each write screen to live data only after it passes there.
Give the agent its own ERP user and licence
Give the agent a user of its own, with rights to the screens its job needs and no more. OWASP, a non-profit security foundation, names the opposite risk excessive agency and traces it to an AI tool with more functions, permissions or independence than its job needs. Its fixes include minimal permissions, enforced in the system the agent writes to, so the ERP refuses what the agent should not do whatever the model decides. If your ERP separates saving an entry from posting it, give the agent the right to save only, and let a person post.
That user may need a licence of its own. Microsoft’s licensing guidance on multiplexing, its term for reaching a product through pooled connections or automation, says multiplexing does not reduce the number of licences its products need. Check your ERP vendor’s terms for an automated user before you start.
Keep passwords in a vault
Anthropic’s security guidance for computer use advises against giving the model access to sensitive data such as login details. Have the control program fetch the ERP password from a password vault and sign in before it hands the screen to the model, and keep the password out of the model’s instructions. Copilot Studio works this way: it keeps computer use passwords in Power Platform storage or an Azure Key Vault you provide, and fills sign-in prompts itself. The account settings to check with each AI provider are in secure AI at work.
A person approves every save
Anthropic, OpenAI and Google each ask for a person to confirm actions with real consequences. When several actions arrive in one reply (a batch), OpenAI asks you to stop before the first one that needs confirmation, and Anthropic asks for the check before each action runs. Gemini can mark an action as needing confirmation. Copilot Studio can email a named reviewer when it detects possibly harmful instructions.
In an ERP, the save button is the natural checkpoint. Put the approval on the data, as in the worked example, and have the control program refuse any save whose screen does not match what the person approved.
Read back every value and block duplicates
Anthropic notes that Claude sometimes assumes an action worked without checking, and suggests telling it to take a screenshot after each step and confirm the result. OpenAI asks you to check the actual outcome of every run as well as the model’s final answer. Do both, and add a check that does not depend on the model: after every save, the control program reads the record back through the read-only database login and compares each field with the approved draft.
Retries create a second risk. If a run fails halfway and starts again, it can enter the same order twice. Before creating any record, check a unique reference such as the customer’s PO number or the supplier’s invoice number. UiPath’s Orchestrator applies the same idea to RPA: its work queues can require each transaction’s reference to be unique, and a duplicate fails the job.
Handle pop-ups, errors and limits
Old ERPs raise warnings, such as a credit hold, a price below cost or a record locked by another user. Write each known warning into the procedure with the right response, and have the agent stop and hand the task to a person on anything it does not recognize. Anthropic notes that dropdowns and scrollbars can be hard to work with the mouse, and suggests keyboard shortcuts instead. Its sample loop stops at a maximum number of iterations, and OpenAI asks for step, time or cost limits on every run.
Keep a screenshot log
Google recommends logging prompts, screenshots, the model’s proposed actions and every action executed, for debugging, auditing and incident response. File the log under the business reference, such as the PO number, so anyone can replay what happened to one order. The screenshots show customer names and prices, so give the log the same access controls as the ERP and a fixed retention period.
Prompt injection from the screen
Prompt injection is text the AI reads that changes what it does. Anthropic warns that Claude may follow instructions it finds in content, including text on web pages and inside images, even when they conflict with yours. On an ERP, that text can come from outside the company: a note a customer typed into your web portal, or a line from a supplier’s PDF that someone copied into a comment field.
Anthropic scans computer use screenshots with classifiers, separate models that flag possible injections, and Google offers injection detection that you switch on. Anthropic adds that its other precautions remain important with the classifiers in place. What limits the damage is the setup around the model. The machine reaches only the servers its job needs, and the agent’s ERP user has limited rights and saves nothing a person has not approved. For the wider picture, see AI security for smaller manufacturers and distributors.
Where the screenshots go
Every screenshot the agent sends shows what is on the ERP screen, such as customer names, addresses, prices and costs. Where that image is processed depends on the model you choose. Processing, also called inference, is the step where the model reads your request and writes its reply.
- Anthropic’s API. Its inference location setting accepts only “global” or “us.” Computer use is eligible for zero data retention, an arrangement agreed with the provider under which it does not store your requests or its answers after the answer is returned. The exception is the models Anthropic calls Covered Models, Claude Fable 5 and 5.1 and Claude Mythos 5 and 5.1, which require 30-day retention.
- Claude on Amazon Bedrock. Anthropic lists computer use on Bedrock as a beta. For Claude Opus 5.5, AWS’s model card lists no in-region option in its Canada (Central) region, and says its US cross-region option, which routes each request across the AWS regions in that group, “keeps data within US and Canada regions,” so a request may be processed in the United States.
- OpenAI. OpenAI’s data controls page lists a Canada region that stores data in Canada, with regional processing marked as unavailable there. In a region without regional processing, OpenAI says it may process and temporarily store data outside the region.
- An open-weight model on your own server. Holo3-35B-A3B and UI-TARS-1.5-7B run on hardware you control, so the screenshots stay on that hardware. Holo3-35B-A3B’s best run on the OSWorld-Verified leaderboard described below scores 82.56%.
Canadian privacy law keeps your company responsible for personal information it sends to a provider. The Privacy Commissioner’s guidelines for processing personal data across borders say the company that transfers personal information stays accountable for it in the hands of its service provider, and should tell customers their information may be processed in another country. In Quebec, section 17 of the private-sector privacy act requires a privacy impact assessment and a written agreement before personal information leaves Quebec. This page is not legal advice. How to get zero data retention in writing is covered in secure AI at work.
Speed, reliability and cost today
Computer use agents have improved quickly. On long workflows they still fail more often than they succeed, and they are slower than a person doing the same work. The published figures:
- Short tasks. OSWorld is a benchmark of 369 tasks in real desktop and web applications. In the original paper, people completed more than 72% of them and the best model 12.24%. On the OSWorld-Verified leaderboard, the updated version on the same site, Claude Fable 5 scores 85.96% and Claude Opus 5 scores 83.39% with a budget of 100 steps. The applications tested include Chrome, LibreOffice, Thunderbird and VS Code. No ERP is in the set.
- Long workflows. OSWorld 2.0 has 108 workflows that take a person a median of about 1.6 hours. On its leaderboard, updated September 17, 2026, the best official run on the full set of task release v2.1, Claude Opus 5 at maximum effort with 500 steps, fully completed 44.33% of the workflows. The authors found that agents lose track of constraints, guess instead of asking the user, and skip verification.
- Speed. Anthropic lists latency as a limitation: computer use can be too slow compared with a person doing the same work. a16z reports that a task a person finishes in two to three minutes can take an agent eight to ten.
- Cost. a16z puts running an agent at roughly $6 to $8 per hour of inference, based on a founder’s estimate, and anywhere from $3 to $15 in practice, depending on how the control program is built.
Our own rough method, for one step of an order entry: models read and bill text in tokens, and by Anthropic’s image formula a 1280 by 720 screenshot works out to 1,196 input tokens. The computer use toolset adds about 4,500 tokens to each request. A step with about three screenshots (Anthropic suggests pruning back to the last three), a page of instructions and a few hundred output tokens is billed at the model’s rates per million input and output tokens, which Claude API pricing lists for each Claude model. Caching the fixed part of each request, which means the provider reuses it from earlier requests at a lower price, lowers the cost of every step, and an order costs as many steps as it takes. We found no provider that publishes seconds per step, so measure the time and the cost of a step on your own screens.
No published benchmark includes an ERP, so the numbers that decide whether to go live come from your own pilot: the share of transactions that finish correctly with no help, the minutes and model cost per transaction, and how often a person had to step in.
How ThriveAI helps
ThriveAI is an AI engineering company in Ottawa. It builds private AI systems on your own data for manufacturers and distributors in Ontario and Quebec, and it works hands on with your team, on your own systems. Derik Lawlis, the founder, leads every project and stays close to the build.
On an old ERP, the work starts with the route table above. Reads move to a read-only database login, and writes use the ERP’s import tools where it has them. A computer use agent takes only the screens that are left, on a machine inside your network with its own ERP user. Nothing is saved in your ERP until the person responsible approves it, and every saved entry is read back and checked.
The platform ThriveAI builds on is designed to keep your data on its own server in Canada. You choose the model that reads it: one that runs on that server, or a hosted model under a written zero data retention agreement. A hosted model may process requests outside Canada, so the contract names the model and its service tier.
Working sessions run on site with your team, in French or English, and ThriveAI also runs hands-on AI training on your own documents. Stage one is a working prototype for one job at a fixed price, and you keep it. For the company and how a project runs, see About ThriveAI.