1. Write down the decision
Begin with a task that somebody can explain without naming a product. “Prepare a draft reply using our returns policy” is more useful than “introduce a chatbot”. Record the incoming information, the expected output, the person responsible and the action that follows. Distinguish producing information from making a consequential decision.
Collect examples of normal work and difficult exceptions. Remove unnecessary personal data before sharing them. A discovery engagement can use these examples to describe scope, identify missing inputs and separate tasks that need judgement from tasks that follow a fixed rule. Include the existing process so the proposed system has something meaningful to be compared against.
2. Choose the simplest suitable method
Rules-based software follows explicit conditions. It is a sensible starting point when the decision can be expressed clearly and the inputs are stable. Predictive machine learning learns patterns from examples and can support classification or forecasting, but depends on suitable training and evaluation data. Generative AI produces new text or other content; fluent output is not evidence that the content is correct.
A language model may help interpret varied messages, while ordinary code checks identifiers and updates records. These approaches can be combined. Compare them on the task itself: how errors appear, whether a result can be explained, how changes are maintained and what happens when information is missing. Do not use a model merely to replace a dependable calculation.
3. Check data access and responsibilities
Make an inventory of the documents, databases and applications involved. Identify who owns each source, how access is granted and whether the information is sufficiently complete. UK organisations should consider UK GDPR and the Data Protection Act 2018 when personal data is involved. A lawful basis, purpose and appropriate safeguards need attention before implementation, not after a demonstration.
For a hosted model, examine contractual terms covering processing, retention, training use and international transfers. For a locally hosted model, account for infrastructure administration and security updates. Local hosting does not remove data protection duties. Discovery should record unresolved questions and assign them to an owner rather than assuming that a technical setting answers a legal question.
4. Define a pilot that can fail safely
A pilot is a bounded implementation used to explore whether the proposed approach works for the task. Keep its access limited, use a clear stopping condition and decide which outputs require approval. A draft-only pilot is often easier to assess than one permitted to send messages or alter operational records. That restriction should be deliberate and visible to users.
Agree acceptance criteria before reviewing results. Depending on the task, assess factual accuracy, missed information, handling time, usability and the effort needed to correct mistakes. Record the severity of errors as well as their frequency. A plausible answer with the wrong contractual condition matters differently from a formatting defect. Avoid a single average score that hides important failure types.
5. Compare operating costs and dependencies
Include integration, support, monitoring and human review in the cost discussion. Model usage is only one component. Long documents can increase processing demands; frequent source changes can create maintenance work. Check provider pricing and contractual limits directly when preparing an estimate rather than relying on a quoted figure that may have changed.
Map dependencies on model providers, search services and business systems. Decide how prompts, configuration and evaluation examples will be recorded so another person can maintain the service. Where possible, keep business rules separate from provider-specific code. Portability involves preserving the behaviour and controls, not simply swapping one API address for another.
6. Leave with a decision, not a slogan
The useful output is a scoped recommendation: proceed with a defined pilot, improve the underlying data first, use conventional automation, or stop. It should state assumptions, exclusions, decision owners and the evidence still needed. A recommendation against an AI build can be the appropriate result when simpler software addresses the problem.
The NIST AI Risk Management Framework provides a reference for identifying and managing AI risks. The ICO’s AI guidance provides a UK data protection starting point. Use these alongside organisation-specific requirements; neither replaces a clear project scope or qualified advice where required.
