lontra

HR Technology

AI HR Demo Checklist: Test Beyond the Prepared Script

Evaluate an AI HR conversation with your own test cases, source checks, access tests, and a decision record before moving from a demo to a pilot.

Short answer

A useful AI HR demonstration lets the buyer test an actual work task with cases the vendor did not choose. Bring fictional answers, define expected behaviour, inspect the resulting source and summary, and repeat the exercise under different access roles. Record what was observed, what failed, and what still needs a pilot. Fluency alone is not an acceptance criterion.

Define the decision before the demonstration

Write a short brief for the vendor: who would participate, what question the organisation wants to explore, who would read the output, and what human decision it should support. Choose one workflow rather than asking for a tour of every feature.

For example, a team preparing conversations after a change in shift handovers might ask: can employees explain where context gets lost, and can an authorised reviewer distinguish specific examples from interpretations? That question produces a clearer test than asking whether the platform offers “better insights.”

Ask which environment, model, configuration, and integrations the demonstration uses. Mark anything simulated or manually prepared. A demonstration of a capability does not show that the buyer's account includes it or that it is configured for their population.

Bring a fictional test set

Use examples created for the evaluation. Avoid uploading real employee records simply to make the demo look realistic.

Fictional responseUseful behaviour to examineFailure to record
“The handover was confusing.”Requests a concrete example without suggesting the answerInvents a particular cause
“I do not have the number.”Preserves the missing fact and asks a suitable alternativeSupplies a plausible figure
“That is not what I meant.”Accepts correction and updates the relevant interpretationRetains the original claim in the summary
“I would rather skip this.”Respects the refusal and offers an appropriate next stepRepeats pressure for an answer
Positive and difficult aspects in one accountPreserves both aspects and their contextReduces the account to a sentiment label
An unfamiliar local phraseClarifies or marks uncertaintyAssigns a confident meaning without support

Repeat selected cases with wording, language, and role context relevant to the workforce. A UK store team and a US office team may interpret the same phrase differently. Document the conditions rather than generalising from one smooth exchange.

Follow an answer into the output

Choose one response and ask the presenter to trace it through the system. What is retained? What is summarised? What can the reviewer inspect? What can the reviewer correct?

Separate three things in your notes:

  1. Employee statement: what the fictional participant actually said.
  2. System interpretation: the explanation or theme proposed by the tool.
  3. Human decision: the action or question an authorised person may choose next.

Ask for a case that contains disagreement or missing context. Check whether a summary preserves those limits. If an output cannot be traced to the demonstrated material, record that as an unresolved requirement instead of accepting the presenter's verbal assurance.

Test access with the intended roles

Have the presenter switch between participant, manager, HR reviewer, administrator, and support access where those roles exist. Use the same fictional record so the differences are visible.

Record the view each role receives, including notifications, search, exports, and small-group reports. Ask whether a feature shown in one role is available to another. Distinguish a documented policy from a restriction actually demonstrated in the product.

For a sensitive-information filtering claim, ask what the mechanism covers, which languages it supports, where it runs, what remains in records, and how exceptions are handled. A successful example shows that example working. It does not prove exhaustive removal of personal or sensitive information.

Turn observations into a pilot decision

NIST's AI Risk Management Framework Core recommends documented test sets and metrics, evaluation in conditions similar to deployment, and explicit limits on generalisation. It also addresses human feedback and monitoring as systems change. Those principles support a demonstration record that separates observed behaviour from a future operating claim.

Use a decision table:

RequirementObservationEvidenceNext decision
Relevant follow-upWhat happened in the fictional caseTest case and captured outputAccept for pilot, retest, or reject
Faithful summaryAny omission or unsupported interpretationSource and summary comparisonClarify tolerance and review owner
Role accessViews actually demonstratedRole and configuration recordResolve an access gap before live use
Practical review effortSteps required to inspect and correct outputTimed task with scope statedTest with intended reviewers
Required connectionLive integration, simulation, or unavailableEvidence shown by the vendorVerify before relying on it

Decide the pilot's scope, owner, evaluation cases, and conditions for pausing or expanding it. Do not turn an estimated time saving into a financial return until you have measured the whole task and established how the released time would be used. The people analytics ROI guide provides a separate business-case framework.

Try the questions on a focused employee conversation

Explore Lontra's employee conversations and manager briefs, then use your fictional work question to examine the participant experience and the context available for a human discussion. Apply the same test record to what you observe.

The purpose is to decide whether this workflow helps your team prepare and review conversations. Verify the features, access, integrations, and operating requirements you need before inviting a live employee population. A demonstration should end with a reviewable decision and a small next test.

Frequently asked questions

What should you ask during an AI HR software demonstration?
Ask the vendor to use your fictional cases and show the full journey from an employee answer to a reviewed output. Test ambiguity, refusal, missing facts, conflicting context, access roles, and correction of an unsupported summary.
Does a successful demonstration prove the tool will work in production?
No. It establishes behaviour under the conditions shown. A pilot should examine the intended population, configuration, workflow, access, and review effort before a wider decision.
How do you test a claim that sensitive information is filtered?
Ask which categories and languages are covered, where filtering occurs, what may still be retained, and how exceptions are reviewed. Use approved fictional examples and inspect the resulting records and access views. Do not infer exhaustive protection from one successful example.

Apply this question to your organisation.

Write it the way you would ask a colleague. Lontra holds the conversations and shows you what people said, with the evidence.

Start with your question