Start with the decision the organisation needs to improve, not the tool a vendor wants to sell. Write down the intended users, the data involved, the human decision that remains accountable and the measurable outcome. If the problem cannot be stated without naming a product, the team is not yet ready to compare options.
Classify the data before the demo
Map personal data, special-category information, donor records, case notes, location data and confidential programme material before anyone uploads a sample. Ask whether the provider uses prompts or outputs for training, where data is processed, how long it is retained and whether administrators can remove it. Use synthetic or de-identified material during early testing.
Test the failure modes
Accuracy is only one risk. Test whether the system invents sources, treats languages differently, exposes confidential text, produces inaccessible output or makes recommendations that staff cannot explain. Record the test cases and keep the results with the procurement decision.
Keep a person accountable
Assign an owner who can stop the pilot. Define which decisions always require human review, how affected people can challenge an outcome and what evidence must be retained. High-impact uses involving safeguarding, eligibility, health, legal status or vulnerable people need heightened review and may be unsuitable for automation.
Negotiate the exit before entry
Require export formats, deletion commitments, incident notification, support expectations and a documented exit route. Calculate staff time, integration work, security review and training alongside licence fees. A low sticker price can conceal a high switching cost.
Run a bounded pilot
Choose a low-risk workflow, a defined user group and a fixed review date. Compare the result with the current process, including errors, rework and staff confidence. Expand only when the evidence supports it; stop when risk, cost or operational burden exceeds the value.