Lean AI for Engineering Businesses
Seven principles for cutting token burn, duplication and rework out of everyday AI use - and how to prove the saving.
The month has barely started and the limits are already gone
You bought a team licence. Ten seats, maybe twenty. The business is deliberately moving AI closer to the centre of how work gets done - quotes, tenders, technical reports, meeting notes, proposals, service responses.
Then the messages start. Someone has hit their usage limit on the 11th of the month. Someone else is locked out mid-tender. Two people ask whether you can upgrade to the next tier. The finance director asks what changed, because the invoice has moved and nothing measurable has.
Look at what is actually happening in those accounts and the pattern is nearly always the same. Every person is rebuilding the same background information from scratch, in every new chat. Everybody has their own version of the same prompt. The company overview, the product specification and the tone-of-voice document are being pasted in again and again, several times a day, by several different people. A frontier reasoning model is being used to tidy a table. Connectors are switched on and quietly pulling material into every message whether it is needed or not. And when one paragraph of a ten-page document is wrong, the whole document gets regenerated.
That is not heavy usage. That is waste - and it is the same waste engineers have been designing out of production processes for a century.
Most businesses respond by trying to make their employees better at prompting. That is solving the wrong problem. In manufacturing, persistent waste is rarely corrected by asking each operator to improvise a better technique. Engineers examine the whole system - inputs, setup, standard work, material flow, quality checks, maintenance and feedback. AI deserves the same treatment.
This article gives you seven practical principles for economising your usage and running AI leaner, plus a way to prove the saving in your own business rather than quoting someone else's percentage.
The adoption curve outran the operating discipline
The numbers are not subtle. In 2024, 78% of surveyed organisations reported using AI in at least one business function, up from 55% the year before. By McKinsey's most recent global survey of AI adoption that figure had reached 88%, with 72% using generative AI specifically.
Access, then, is close to universal. Value is not. Only 39% of respondents attribute any level of EBIT impact to AI at the enterprise level, and most of those put it below 5%. Roughly six in a hundred organisations qualify as genuine high performers.
EBIT impact, in plain English: EBIT means earnings before interest and tax - your operating profit before the bank and HMRC take their share. "EBIT impact" simply means the AI work showed up in the profit line, not just in how busy or clever it felt. Fewer than four in ten organisations can say that at all.
The reason is not the model. High performers and laggards are largely using the same tools. What separates them is whether the work around the tool changed. When McKinsey tested 25 organisational attributes against reported EBIT impact, the redesign of workflows had the biggest effect of any of them - and only 21% of organisations using generative AI said they had fundamentally redesigned even some workflows. The most effective lever available was also the one being pulled least.
For an engineering business, the consequence is familiar. You have introduced a capable machine into a process that was never designed around it. Output varies by operator. Setup is repeated for every job. Nobody can tell you the cost per approved deliverable, because nobody baselined it. And because AI makes producing output cheap, the waste does not announce itself as waste - it arrives disguised as productivity.
Scaling an inefficient AI practice does not remove the waste. It industrialises it.
Operational Overview
Prompting is not the bottleneck - The largest avoidable losses come from system design, not from individual phrasing.
Redesign is the differentiator - Of 25 organisational attributes McKinsey tested, workflow redesign had the biggest effect on EBIT impact, yet only 21% of adopters had done it.
AI creates recognisable Lean waste - Overproduction, over-processing, inventory, defects and duplication all have direct equivalents in how teams use AI.
Repeated context has a real cost - Prompt caching exists precisely because sending the same material over and over is expensive to process.
Right-size the machine - A frontier reasoning model on a formatting task, and connectors left permanently switched on, are both over-processing.
Standardise with skills, not a prompt folder - Reusable skills take inputs, so the team stops copying, pasting and re-editing the same instructions.
Start with a baseline, not a benchmark - The saving is workflow-dependent. Measure one repeatable process before and after, and stop quoting other people's percentages.
Quick Links
The terms, in plain English The eight wastes, translated for AI
- Principle 1: Supply context once, not every time
- Principle 2: Match the model to the task
- Principle 3: Build controlled project knowledge, not infinite chats
- Principle 4: Make knowledge modular and retrievable
- Principle 5: Build skills, not a prompt library
- Principle 6: Modify components, do not rebuild assemblies
- Principle 7: Govern lightly, measure deliberately The sustainability question, answered honestly Where to start
What does Lean AI actually mean?
Lean AI applies established Lean and engineering-management principles to how a business runs generative AI. Rather than optimising individual prompts, it designs the surrounding system - approved context, right-sized models, reusable knowledge, standard work instructions, controlled revisions and measurement - so that repeatable tasks produce consistent output with less duplicated effort and rework.
The terms, in plain English
You do not need to be technical to run this, but a handful of words do the heavy lifting. Here they are without the jargon.
Token
The unit AI platforms count. Roughly three-quarters of a word. Everything you send and everything the model sends back is measured in tokens, and that is what your plan limits and your bill are based on.
Context window
The model's working memory for a single conversation - how much it can hold in view at once. Fill it with material the task does not need and you pay for it twice, in cost and in quality.
Usage limit
The cap on your seat or plan. Hitting it mid-tender is the visible symptom of an invisible process problem.
Prompt caching
A platform feature that stores stable, repeated material so it is not reprocessed from scratch every time. Cheaper and faster on the repeats.
Connector
A live link between your AI tool and another system - your drive, your inbox, your CRM. Useful when the job needs it. Costly when it is left on by default.
Skill
A saved, reusable set of instructions that accepts inputs. Closer to a setup sheet than a saved paragraph of text - you feed it the job details and it applies the standard method.
Retrieval
Pulling only the relevant section of your knowledge into the conversation, rather than the whole document library.
Agent
AI that carries out a sequence of steps on its own rather than answering one question at a time.
The nine wastes, translated for AI
Lean gives us a vocabulary for waste that engineering teams already know. It maps onto AI usage almost directly.
- Overproduction - generating a full report when one field was needed
- Waiting - people queuing on long generations or hunting for approved sources
- Motion - shuttling information manually between chats, documents and systems
- Over-processing - excessive context, or a frontier model on a trivial task
- Inventory - obsolete chats, duplicate documents, uncontrolled prompt versions
- Defects - hallucinations, unsupported claims, inconsistent terminology, rework
- Underused talent - experts rebuilding prompts and correcting preventable errors
- Duplication - several people independently rebuilding the same setup
- Validation - this is one CREATED by AI, we need to check facts, citations and calculations.
That second last one is worth dwelling on, because it is not new. APQC's research into knowledge-work productivity found the average knowledge worker spending
8.2 hours per week searching for, recreating and duplicating information - 2.8 hours searching or requesting, 2.0 recreating work that already existed, 1.7 supplying duplicate answers and 1.7 tracking down the right person. That study surveyed 982 knowledge workers and predates generative AI entirely.
Which is precisely the point. Unstructured AI adoption does not create the duplication problem. It reproduces an existing one in a new interface, faster and at greater volume.
The last one is also worth a mention - it’s new, you could akin it quality management, but this is hidden and not the obvious defect like on a production line. Validating and checking critical facts - can sometimes be longer than generation if not managed well and could be the biggest wastes of all. But if managed well with tighter prompts, second AI checking it and flagging risks, this can be minimised
Engineer's Note:
The first thing I ask when a technical business tells me AI is not delivering is to show me two people doing the same task. Nine times out of ten they are using different prompts, different source documents and different terminology, and both are rebuilding the setup from nothing. That is not a model problem. That is a standard work problem, and we have known how to fix those for a century.
Principle 1: Supply context once, not every time
There is a common assumption that shorter prompts are the goal. They are not. The goal is to stop paying repeatedly to process the same stable material.
Prompt caching is the clearest evidence that repeated context carries a real technical and economic cost. OpenAI's prompt caching documentation states that caching can reduce time-to-first-token latency by up to 80% and input token costs by up to 90%. Anthropic and Google Cloud offer comparable mechanisms.
The mechanism matters more than the headline. Caching works on stable prefixes, which means your standardised system instructions and approved reference material should sit at the front of the prompt, with the variable task data behind them. Get the order wrong and the cache never hits.
Two honest qualifications. The percentages are "up to" figures, dependent on workload, model and cache-hit rate. And caching is a cost and latency reduction, not a removal of tokens from the model's context window - the material is still being reasoned over, it is simply not being recomputed from scratch.
The analogy an engineer will recognise: it is the difference between setting a machine up once for a repeat production run and performing the full setup for every single component.
In one line for the MD: Your team is paying to re-explain the same company background several times a day.
Practical steps
- Write one approved company and product briefing document
- Store it where the AI tool can reach it, not in someone's downloads folder
- Put stable instructions first in the prompt, job-specific detail last
- Ban pasting the same background in fresh every time
- Give one person ownership of keeping that document current
Principle 2: Match the model to the task
Using the largest reasoning model for every job is like reaching for a two-foot adjustable spanner to wind a clock. It will not do the job better. It will do it slower, at greater cost, and with a real chance of damage.
Most platforms now offer a range - fast, light models for routine work and slower reasoning models that think through a problem before answering. Reasoning models earn their cost on genuine analysis: interpreting a specification, stress-testing a tender assumption, working through a technical argument. They do not earn it on reformatting a table, tidying meeting notes or drafting a short internal email.
Connectors follow exactly the same logic. A live link to your drive, inbox or CRM is valuable when the task needs it. Left switched on permanently, every connector adds its tooling and whatever it retrieves into the conversation - on every message, whether the job called for it or not. It is the equivalent of running every machine on the line because one of them is in use.
In one line for the MD: You are running the biggest machine in the shop to cut a bracket, with everything else idling alongside it.
Practical steps
- Set a default model for routine work and reserve the reasoning model
- Write a simple rule of thumb: analysis and judgement get the big model, formatting and drafting do not
- Turn connectors on for the job, then turn them off again
- Review which connectors are enabled by default across the team this month
- Spot-check one week of usage and ask which tasks actually needed the heavy model
Principle 3: Build controlled project knowledge, not infinite chats
The instinct when a conversation is working well is to keep it going forever. This is the equivalent of never closing out a job file.
- Use a focused conversation during an active work phase
- Record approved terminology, constraints and decisions as they are settled
- Move stable information into a maintained project file or knowledge base
- Remove superseded information rather than letting it accumulate
- Start a fresh, phase-specific conversation when the task changes materially
Retain the approved specification and the change record. Not every informal conversation that preceded them.
This is the step most businesses skip, and it is the one that compounds. We wrote about the wider pattern in our piece on AI adoption in business, which covers why so many organisations jump from casual experimentation straight to agents without building the layer in between.
In one line for the MD: A chat that never ends is a job file that never gets closed, and you are paying to carry it.
Practical steps
- Close the chat when the phase ends and start a new one
- Move anything settled and reusable into the project file
- Delete superseded drafts rather than leaving them in the thread
- Keep the approved output and the change record, not the working conversation
- Name the person responsible for tidying each project space
Principle 4: Make knowledge modular and retrievable
Retrieval quality depends on how knowledge is segmented. If your reference material is one enormous document, retrieval pulls in irrelevant content. If it is shredded into fragments, retrieval loses the context needed to interpret a fact safely.
The principle to work to: use the smallest complete unit of knowledge, not the smallest possible file.
Structure matters here too. It is easy to let project folders in your chat platform grow, because adding one more document always feels harmless. Several tightly scoped projects pull in far less context than one sprawling one - even with caching in place.
For an engineering business, the natural modules are the ones you already maintain - product specifications, approved technical terminology, design-review procedures, bid qualification criteria, risk and assumptions registers, change-control rules, commissioning checklists, case-study evidence, quality and compliance requirements.
In one line for the MD: One giant folder makes every task expensive. Several tidy ones do not.
Practical steps
- Split one oversized project space into two or three scoped ones
- Keep each document to a single complete subject
- Remove superseded versions rather than adding new ones alongside
- Name files so a colleague can find the right one without opening three
- Review project contents quarterly, as you would a controlled document set
Does file format affect AI performance?
Markdown is easier to maintain, version and segment than visually complex formats, because heading hierarchy and section boundaries are explicit. That helps retrieval quality. It does not mean the file extension itself produces a fixed token saving. Word and PDF remain correct where page layout, signatures, tracked changes, diagrams or controlled publication are part of the requirement.
Principle 5: Build skills, not a prompt library
A folder of clever prompts sounds like standardisation. In practice it rarely behaves like it.
Someone opens the document, finds the tender prompt, copies it, pastes it, then edits it for this particular client, this particular scope and this particular deadline. Every use is a fresh edit, and every edit is an opportunity to drop a constraint, change the terminology or lose the acceptance criteria. When the prompt has been written generically enough to suit everyone, it produces output specific to no one - and the rework lands on the reviewer.
A skill fixes that. A skill is a saved, reusable instruction set that takes inputs. Instead of copying a paragraph and rewriting it, the user supplies the variables - client, scope, deadline, document type - and the standard method is applied the same way every time. Update the skill once and everybody is working to the current revision, without a broadcast email asking people to use version 4.
A properly built skill carries: objective, approved sources, required inputs, constraints, output format, acceptance criteria, validation checks, human review point, and version and owner. That is not a magic command. That is a setup sheet with an inspection instruction attached.
The defensible claims here are improved consistency, auditability, less setup effort and less rework at review. We would not attach a universal percentage to it, and neither should anyone else.
In one line for the MD: Copy-paste prompting is uncontrolled work instructions. A skill is a controlled one.
Practical steps
- Pick the one task the team runs most often and build a skill for it
- Define the inputs the user must supply, so nothing gets guessed
- Write the acceptance criteria into the skill, not into people's heads
- Version it and name an owner, exactly as you would a drawing
- Retire the loose prompt document once the skill is live
Scores the reader against the seven principles and routes them to the appropriate next step based on their answers.
Principle 6: Modify components, do not rebuild assemblies
When something is wrong with a generated document, the reflex is to regenerate the whole thing. In engineering terms, that is scrapping an assembly to correct one component.
- Identify the component or section being changed
- State the defect or the required outcome
- List the dependencies and content that must stay unchanged
- Request only the replacement component where practical
- Perform a final interface and consistency check
This reduces generated output and, more importantly, reduces the review burden - because the reviewer is checking one section against known criteria rather than re-reading an entire document for changes they cannot see.
In one line for the MD: Regenerating a ten-page document to fix one paragraph costs twice - once to produce it and once to re-check it.
Practical steps
- Ask for the section, not the document
- Say explicitly what must not change
- Paste back only the paragraph being corrected as context
- Run one consistency check at the end rather than after every edit
- Train the team on this first - it is the fastest visible saving
Principle 7: Govern lightly, measure deliberately
Governance in this context is not a compliance layer bolted on afterwards. It is what makes the rest of the system repeatable.
The lightweight set: approved tools, data-access rules, named ownership of skills and workflows, source traceability, output review thresholds, usage and cost reporting, version control, a defect feedback route and defined escalation points.
McKinsey's evidence supports this being an enabler rather than a brake - governance oversight, well-defined KPIs and workflow redesign all correlate with reported business impact, and fewer than one in five organisations were tracking KPIs for their generative AI solutions at all.
In one line for the MD: You do not need a committee. You need to know who owns the tender skill and what happens when it gets something wrong.
Practical steps
- Name an owner for each standard workflow
- Agree which outputs need a human sign-off before they leave the building
- Pull the usage report monthly and look at it
- Set two or three KPIs per workflow and nothing more
- Give people one route to report a bad output, and act on it
Engineer's Note:
Every time I raise governance with an MD, they hear "committee". What I actually mean is knowing who owns the tender prompt and what happens when it produces something wrong. That is not bureaucracy. That is the same discipline you already apply to a drawing revision.
The sustainability question, answered honestly
There is a temptation to attach an environmental claim to efficiency work. Be careful here.
The system-level trend is real and well documented. The IEA projects global data-centre electricity consumption more than doubling by 2030, to around 945 TWh - slightly more than Japan's entire current electricity consumption, with AI the most important driver of that growth.
What follows from that is a limited claim, and it should stay limited. Avoiding unnecessary processing is directionally consistent with reducing AI's computational burden. Token counts alone cannot produce a reliable carbon figure. Choosing an appropriately sized model, limiting unnecessary output, retrieving only relevant context and caching stable material are all defensible on those grounds. Generic energy-per-prompt or water-per-prompt comparisons are not, and quoting them will cost you credibility with exactly the technical audience you are trying to convince.
Where to start
The prize is worth having. In a controlled experiment with 453 professionals, access to generative AI cut average completion time on midlevel professional writing tasks by 40% and raised assessed output quality by 18%. That study used GPT-3.5, in early 2023, on 20 to 30 minute writing tasks - press releases, grant applications, short reports, sensitive emails. It is not a claim about engineering work generally. It is evidence that when the task is well defined, the gains are substantial and measurable.
Well defined is the operative phrase. Which brings us back to the baseline.
Choose one repeatable workflow - tender preparation, technical-report drafting, meeting-to-action processing. Then measure it honestly, before you change anything: total input and output tokens (or the percentage of the monthly allowance it consumes), AI cost per approved deliverable, human setup time, number of revision cycles, time to approval, percentage approved without major rework, retrieval or citation accuracy, cache-hit rate where available, repeated prompts and duplicated source files.
Redesign the process using the seven principles above. Then measure it again.
The percentage improvement will differ between a tender, a technical report, a service response and a design review. That is exactly why an engineering business should not begin with a universal claim about token savings. It should begin with a baseline. That is the difference between Lean AI as an attractive analogy and Lean AI as an engineering discipline.
Applying this in your business
If you lead a technical or industrial SME, the practical starting point is smaller than it sounds. You do not need a platform decision, a budget round or a transformation programme. You need one workflow that runs often enough to matter, a fortnight of honest measurement, and someone named as its owner.
If the team needs the underlying habits first - right-sizing models, controlling context, building skills instead of pasting prompts - that is what our AI masterclasses and training for technical teams cover at.
The businesses getting value from AI are not the ones with the best prompts. They are the ones that treated AI as a change to the production process and documented it accordingly. That is a capability engineering firms already have - it is simply being applied to a machine they have not yet recognised as part of the line.
FAQ
Is Lean AI just prompt engineering with a different name?
No. Prompt engineering optimises a single instruction. Lean AI designs the system around every instruction - what context is supplied, which model runs it, where knowledge lives, how work is standardised, how revisions are controlled and how the whole thing is measured. Better prompts help. They do not fix a process that regenerates entire documents to correct one paragraph.
Why does our team keep hitting usage limits?
Almost always because the same background material is being reprocessed many times a day by many people, often on a model far larger than the task requires, with connectors pulling in content nobody asked for. Fix the setup and the same licence goes considerably further before you consider a bigger plan.
How much will this actually save us?
We cannot tell you, and you should be suspicious of anyone who can. The saving depends on the workflow, the model, the volume and how much rework you are currently absorbing. That is why the recommendation is to baseline one process rather than adopt a published benchmark.
Do we need to buy a RAG system or enterprise search to do this?
Not to begin. The early principles - supplying context once, right-sizing the model, maintaining controlled project knowledge and structuring it into coherent modules - can be implemented with the tools most businesses already have. Retrieval infrastructure becomes worthwhile when the volume of knowledge exceeds what can sensibly be supplied by hand.
Won't standardising make output generic?
It has the opposite effect in practice. A skill fixes the approved sources, the terminology and the acceptance criteria, then takes the job-specific detail as an input - which is what makes output consistently correct for your business rather than generically plausible. The variation you lose is the variation you did not want.
Where does model choice fit in?
Under over-processing, and it now has its own principle. Using a frontier reasoning model to reformat a table is the AI equivalent of running a five-axis machine to cut a bracket. Match the model to the task, switch connectors off when the job does not need them, and measure the difference rather than assuming it.
How does this relate to agentic AI?
Agents amplify whatever system they are placed into. If your context is uncontrolled, your knowledge is unstructured and your instructions are undocumented, an agent will reproduce those problems autonomously and at speed. The principles here are the groundwork that makes agents safe to deploy, not an alternative to them.
Who should own this in a business our size?
One named person, usually an operations or technical lead rather than IT. The role is not to write prompts for everyone. It is to own the standard work, maintain the knowledge base and hold the measurement. In a business under 200 people this is rarely a full-time job, but it must be somebody's job.
About the author
Stefan Buss BSc Eng (Ind) is the founder of Sales and Marketing Engineers Ltd, working with industrial and technical SMEs on sales and marketing strategy, AI adoption and the systems that connect the two. He has over 20 years in digital marketing and an industrial engineering background, and works with engineering businesses across the UK on building AI into operations rather than bolting it onto admin. You can find him on
LinkedIn at or read more about our
AI services for sales and marketing at.
References
- Stanford HAI, AI Index Report 2025
- McKinsey, The State of AI, November 2025
- McKinsey, The State of AI: How Organizations Are Rewiring to Capture Value, March 2025
- APQC, KM Makes Knowledge Workers More Productive and Less Stressed Out
- OpenAI, Prompt Caching documentation
- Noy and Zhang, Experimental evidence on the productivity effects of generative artificial intelligence, Science, 2023
- International Energy Agency, Energy and AI, 2025












