Overview
Scoring applicant submissions fairly and consistently is hard at volume. Real Events' ChatGPT integration provides a suggested score and a written rationale for each applicant's responses, so your team can review faster and rank consistently. Each response is scored against a configurable rubric, and ChatGPT explains why it gave each score.
ChatGPT scores the subjective, written responses — the questions an applicant answers and any background detail you collect. Objective data you already hold (counts, flags) is best scored by formulas so it stays consistent; ChatGPT is used only where judgment of written text adds value.
The AI score is a suggestion. It ranks and triages applicants so your team spends its time on judgment and verification — it never makes the final decision. A reviewer can override any score.
What gets scored
On each Evaluation record, ChatGPT scores six items, each 0–5, with a short rationale for each:
- Q1–Q5 — five question responses, mapped to whatever five questions your program asks.
- RC — Credibility — the applicant's relevant background and track record, drawn from two credibility input fields.
The responses and background detail come from your own data (a portal, an import, or manual entry) and are stored on the Evaluation record before scoring.
Configuring the rubric
The criteria ChatGPT uses live in a Scoring Rubric record (a Custom Metadata record). The package ships a Default rubric so scoring works immediately on install, with generic criteria.
To use your own criteria, create a new Scoring Rubric record with your program-specific prompt, mark it Active, and deactivate the Default record. Only one rubric should be Active at a time. Because the rubric is configuration rather than code, you can refine it at any time without an upgrade, and each change can be tracked by version.
Each rubric record holds the System Prompt (your criteria and the required output format), the Model, Temperature, Endpoint, an Active flag, and a Version label that is stamped onto every score for auditability.
How to score an Evaluation
Three ways to start scoring, all doing the same thing:
- A single record — set the Evaluation's AI Scoring Status to Ready to Score; it's picked up and scored automatically.
- A group — select multiple Evaluations from a list view and use the Score with AI action.
- Automatically — a scheduled process scores any Evaluation marked Ready to Score on a regular basis, so new submissions are scored as they arrive.
To see results, refresh the page. There may be a delay from ChatGPT; the status updates on its own once the response is received.
Reading the results
When scoring finishes, the AI Scoring Status changes to one of:
- Scored — complete; score and rationale fields are populated for Q1–Q5 and RC.
- Verify Claims — complete, and the credibility score was high. This prompts a human to verify the applicant's claimed background before the score counts; the rationale field shows what was claimed, so it's easy to spot-check.
- Failed — ChatGPT could not be reached or the response could not be read. Set the status back to Ready to Score to retry. On a new install, a "Failed" result usually means a setup step was missed — see Installation & Setup.
- Manual Override — indicates a reviewer changed a score by hand. The original AI score and rationale remain visible for context.
Each scored Evaluation also records, for the audit trail, the AI Model used, the rubric version, and when it was scored — so you always know which criteria scored each record and any decision can be explained later.
Guiding principles built into scoring
- Scores judge the substance of a response — specificity, credibility, clarity — and never reward or penalize appearance, gender, or other personal characteristics.
- Credibility scores reflect what an applicant claimed, not verified fact; high scores are flagged for human verification.
- Most genuine responses score in the 2–4 range; a 5 is reserved for specific, credible, clearly differentiated answers. A rubric that scores everything high isn't doing its job — the point is to separate strong submissions from generic ones.
Tips
- Calibrate before it counts. Score a varied set of real submissions and compare the AI's scores to your team's judgment. Where they differ, refine the rubric wording. Treat early scores as tuning, not final input.
- Freeze the rubric per cycle. Tune freely during calibration, then keep the rubric stable for a selection round so all applicants are measured on the same standard. If you must change it mid-cycle, re-score everyone so rankings stay comparable.
Was this helpful?
Last updated 1 month ago