An AI exit interview can help ask relevant follow-up questions and organize qualitative feedback. It still needs a defined purpose, representative testing, source traceability, privacy controls, and human review. Treat collection, coding, and decision support as separate capabilities rather than one promise of intelligence.
Start with the job, not the AI label
"AI exit interview" can describe several different functions:
- inviting a departing employee and presenting questions;
- choosing a follow-up based on the previous answer;
- transcribing speech;
- translating a response;
- proposing codes or themes;
- drafting a summary;
- suggesting questions for a human reviewer.
Each function has a different failure mode. A fluent follow-up can lead the participant. A transcript can omit a qualification. A translation can flatten local meaning. A plausible summary can attribute an interpretation to the employee that the source never supported.
Write an intended-use statement for each function before comparing products. For example: "The system may propose a follow-up that helps a participant clarify a voluntary response. It may not diagnose intent, assess truthfulness, or score the employee or manager."
That statement gives a procurement team something testable. It also keeps an appealing demonstration from quietly expanding the use case.
Test the interview experience separately
Build a small set of fictional responses that reflects the people and conditions in scope. Include:
- a short, clear answer;
- an ambiguous answer such as "growth was limited";
- a response that changes the subject;
- disagreement with the question's premise;
- mixed positive and negative experience;
- a term with local or occupational meaning;
- a participant who declines to elaborate;
- a possible safety, discrimination, or health disclosure;
- at least one language other than the program's main language;
- an accessibility case relevant to the workforce.
For each response, decide in advance what a suitable follow-up should accomplish. It might ask for an example, clarify a term, or give the participant a chance to correct the interpretation. It should also respect a refusal and avoid asking for unnecessary personal information.
Use a worksheet like this:
| Test case | Expected behavior | Failure to record |
|---|---|---|
| "I could not progress" | Ask what progression meant in this context | Assumes promotion was the issue |
| "I would rather not name anyone" | Continue without requiring a name | Pressures participant to identify a person |
| Mixed account | Preserve both helpful and difficult aspects | Summarizes only the dominant sentiment |
| Unfamiliar local phrase | Ask for clarification or flag uncertainty | Invents a confident interpretation |
| Sensitive disclosure | Follow the approved escalation wording | Promises secrecy the process cannot provide |
Measure more than completion. Review whether the follow-up was relevant, neutral, proportionate, and faithful to what the person said. Record where a human reviewer disagrees with the system.
Test analysis against source passages
An AI-generated theme is useful only when a reviewer can inspect the evidence behind it. Create a second test set for analysis with known passages and expected coding decisions. Include overlapping themes, contradictory responses, missing answers, repeated comments from one person, and a cohort too small to report separately.
Check whether the system:
- keeps direct statements separate from analyst interpretation;
- maps every theme to the relevant source passages;
- preserves qualifications, uncertainty, and disagreement;
- reports participant and response denominators;
- avoids counting several passages from one person as several people;
- shows uncoded or low-confidence material for review;
- lets an authorized reviewer correct a code or summary;
- records the correction and the model or rule version used.
For a manual codebook and worked example, use the exit interview analysis framework.
Put AI boundaries into the operating procedure
The NIST AI Risk Management Framework recommends documenting scope, knowledge limits, human roles, test methods, representativeness, uncertainty, and performance under conditions similar to use. The NIST AI RMF Core is a useful structure for vendor review and internal ownership.
For an exit program, define at least these boundaries:
- AI output is a proposal for review, not a finding about a person.
- A theme count does not establish a cause of turnover.
- The system does not infer protected characteristics, mental state, honesty, future behavior, or performance.
- An individual or manager is not ranked from exit feedback.
- No employment action starts automatically from an AI summary.
- Exceptional disclosures follow a named human process.
- Reviewers can inspect the source, document disagreement, and stop use of a failing feature.
In the United States, the EEOC states that using AI in employment decisions must comply with federal civil-rights law. Its AI and algorithmic fairness initiative and AI and ADA resources are relevant when a proposed workflow could affect applicants or employees. Exit comments should not become an unexamined input to another employment system.
Run a fictional side-by-side review
Suppose 12 departing employees are invited, eight participate, and six answer a question about development. Three describe unclear progression, one describes limited training time, one reports a positive internal move, and one gives no example.
System A summarizes: "Development problems caused half of departures."
System B reports: "Three of six respondents to the development question described unclear progression. The interviews do not establish why all 12 employees left or whether progression caused the departures." It links the three passages and shows the different training-time comment separately.
System B gives the reviewer a more defensible starting point. The distinction comes from denominators, source links, and restrained interpretation, not from a more confident writing style.
Repeat this review with a different cohort, language, job family, and interviewer configuration. A single clean demonstration is not evidence that the system performs consistently in your environment.
Review privacy, access, and retention
AI does not make an exit interview anonymous. A manager may identify a speaker from role, timing, location, or a distinctive incident even when a name is removed.
The UK's Information Commissioner's Office says organizations should explain why worker information is collected, the lawful basis, retention period, recipients, rights, and other key uses. It also emphasizes fair, lawful, transparent processing and collecting only justified records. See the ICO guidance on employment records.
Ask the vendor to demonstrate access for each role. Then inspect whether the same protections apply to search, notifications, filters, exports, support access, logs, and backups. Document what happens when a participant raises a matter the employer may need to escalate. The companion guide to confidential exit interviews provides an access and notice worksheet.
Decide with a recorded pilot
A useful pilot has written success and stop conditions. Record:
- who was eligible and who participated;
- which languages, devices, and accessibility needs were tested;
- which AI functions were enabled;
- the approved question and follow-up boundaries;
- analysis agreement and disagreement against the test set;
- unsupported inferences and important omissions;
- participant and reviewer feedback;
- privacy, security, legal, and employee-relations review;
- the owner and date for each unresolved risk.
Do not accept a single "accuracy" percentage without the task, dataset, denominator, reviewer standard, and error distribution. A transcription score does not validate thematic coding. Agreement on common themes does not validate rare or sensitive cases.
Where Lontra may fit
Lontra is an employee conversation platform, not a dedicated exit-management suite, HR system of record, or automated employment decision tool. A team can use one focused campaign, up to 30 invitations over 60 days with no credit card, to evaluate the conversation, review, and manager-brief experience.
Review the current platform capabilities, then run the same fictional test set and decision record described above. Verify any required export or HR-system integration separately. In Lontra's current public model, managers receive a brief rather than raw employee responses, and aggregate HR views use a minimum of five respondents. That threshold is a reporting control, not a guarantee of anonymity.


