Short answer
A useful AI HR demonstration lets the buyer test an actual work task with cases the vendor did not choose. Bring fictional answers, define expected behaviour, inspect the resulting source and summary, and repeat the exercise under different access roles. Record what was observed, what failed, and what still needs a pilot. Fluency alone is not an acceptance criterion.
Define the decision before the demonstration
Write a short brief for the vendor: who would participate, what question the organisation wants to explore, who would read the output, and what human decision it should support. Choose one workflow rather than asking for a tour of every feature.
For example, a team preparing conversations after a change in shift handovers might ask: can employees explain where context gets lost, and can an authorised reviewer distinguish specific examples from interpretations? That question produces a clearer test than asking whether the platform offers “better insights.”
Ask which environment, model, configuration, and integrations the demonstration uses. Mark anything simulated or manually prepared. A demonstration of a capability does not show that the buyer's account includes it or that it is configured for their population.
Bring a fictional test set
Use examples created for the evaluation. Avoid uploading real employee records simply to make the demo look realistic.
| Fictional response | Useful behaviour to examine | Failure to record |
|---|---|---|
| “The handover was confusing.” | Requests a concrete example without suggesting the answer | Invents a particular cause |
| “I do not have the number.” | Preserves the missing fact and asks a suitable alternative | Supplies a plausible figure |
| “That is not what I meant.” | Accepts correction and updates the relevant interpretation | Retains the original claim in the summary |
| “I would rather skip this.” | Respects the refusal and offers an appropriate next step | Repeats pressure for an answer |
| Positive and difficult aspects in one account | Preserves both aspects and their context | Reduces the account to a sentiment label |
| An unfamiliar local phrase | Clarifies or marks uncertainty | Assigns a confident meaning without support |
Repeat selected cases with wording, language, and role context relevant to the workforce. A UK store team and a US office team may interpret the same phrase differently. Document the conditions rather than generalising from one smooth exchange.
Follow an answer into the output
Choose one response and ask the presenter to trace it through the system. What is retained? What is summarised? What can the reviewer inspect? What can the reviewer correct?
Separate three things in your notes:
- Employee statement: what the fictional participant actually said.
- System interpretation: the explanation or theme proposed by the tool.
- Human decision: the action or question an authorised person may choose next.
Ask for a case that contains disagreement or missing context. Check whether a summary preserves those limits. If an output cannot be traced to the demonstrated material, record that as an unresolved requirement instead of accepting the presenter's verbal assurance.
Test access with the intended roles
Have the presenter switch between participant, manager, HR reviewer, administrator, and support access where those roles exist. Use the same fictional record so the differences are visible.
Record the view each role receives, including notifications, search, exports, and small-group reports. Ask whether a feature shown in one role is available to another. Distinguish a documented policy from a restriction actually demonstrated in the product.
For a sensitive-information filtering claim, ask what the mechanism covers, which languages it supports, where it runs, what remains in records, and how exceptions are handled. A successful example shows that example working. It does not prove exhaustive removal of personal or sensitive information.
Turn observations into a pilot decision
NIST's AI Risk Management Framework Core recommends documented test sets and metrics, evaluation in conditions similar to deployment, and explicit limits on generalisation. It also addresses human feedback and monitoring as systems change. Those principles support a demonstration record that separates observed behaviour from a future operating claim.
Use a decision table:
| Requirement | Observation | Evidence | Next decision |
|---|---|---|---|
| Relevant follow-up | What happened in the fictional case | Test case and captured output | Accept for pilot, retest, or reject |
| Faithful summary | Any omission or unsupported interpretation | Source and summary comparison | Clarify tolerance and review owner |
| Role access | Views actually demonstrated | Role and configuration record | Resolve an access gap before live use |
| Practical review effort | Steps required to inspect and correct output | Timed task with scope stated | Test with intended reviewers |
| Required connection | Live integration, simulation, or unavailable | Evidence shown by the vendor | Verify before relying on it |
Decide the pilot's scope, owner, evaluation cases, and conditions for pausing or expanding it. Do not turn an estimated time saving into a financial return until you have measured the whole task and established how the released time would be used. The people analytics ROI guide provides a separate business-case framework.
Try the questions on a focused employee conversation
Explore Lontra's employee conversations and manager briefs, then use your fictional work question to examine the participant experience and the context available for a human discussion. Apply the same test record to what you observe.
The purpose is to decide whether this workflow helps your team prepare and review conversations. Verify the features, access, integrations, and operating requirements you need before inviting a live employee population. A demonstration should end with a reviewable decision and a small next test.
Frequently asked questions
- What should you ask during an AI HR software demonstration?
- Ask the vendor to use your fictional cases and show the full journey from an employee answer to a reviewed output. Test ambiguity, refusal, missing facts, conflicting context, access roles, and correction of an unsupported summary.
- Does a successful demonstration prove the tool will work in production?
- No. It establishes behaviour under the conditions shown. A pilot should examine the intended population, configuration, workflow, access, and review effort before a wider decision.
- How do you test a claim that sensitive information is filtered?
- Ask which categories and languages are covered, where filtering occurs, what may still be retained, and how exceptions are reviewed. Use approved fictional examples and inspect the resulting records and access views. Do not infer exhaustive protection from one successful example.