For HR Business Partners ·
What you'll accomplish
A calibration deck usually means an evening pulling ratings out of one export, comp bands out of another, and last quarter's 9-box out of a third, then reconciling all of it by hand before the meeting. This guide sets up a repeatable ChatGPT workspace that takes those three exports and returns a single prep sheet, grouped by likely performance tier, with the mismatches already flagged (a top rating sitting well under the comp band's midpoint, a "3" rating with three years of no movement) so the calibration conversation starts from a clean draft instead of a blank spreadsheet.
What you'll need
Open your ratings export and your comp export in Sheets or Excel. Replace every employee name with a stable ID (Employee 1, Employee 2, or a code that matches across both files) so you can still link the two exports together without a real name appearing anywhere.
What you should see: Two CSV files where every row is identified by an ID code, with no name column anywhere. Troubleshooting: If your HRIS export includes email addresses or employee numbers tied to a public directory, strip those too. An ID that maps back to a real person defeats the purpose.
Open chatgpt.com and start a new chat. Click the paperclip ("+") icon in the message bar, select your anonymized ratings CSV, then repeat to add the comp CSV to the same conversation.
What you should see: Both file names appear as attachments above the message box before you send anything. Troubleshooting: If the upload option only shows document reading and not code execution, confirm you're on Plus. free accounts get limited file reading without the data analysis features this workflow needs.
Type a prompt asking it to join the two files on the ID column and group employees into likely 9-box tiers based on the rating scale your company uses.
What you should see: ChatGPT responds with a table showing each employee ID, rating, comp band position, and proposed tier, plus a short note on how it built the join. Troubleshooting: If the IDs don't match between files, ChatGPT will usually flag unmatched rows. Check your two source files use the exact same ID for the same person before re-uploading.
Follow up in the same conversation asking it to flag specific patterns worth discussing in calibration.
What you should see: A short, separate list of flagged employee IDs with a one-line reason for each flag.
Before you use this output in a meeting, read through the flags and tiers for language that could read as a stand-in for a protected characteristic (age, disability status, caregiver status, tenure used as a proxy for either).
What you should see: A prep sheet you'd be comfortable defending if someone asked why a particular employee was flagged.
Ask ChatGPT to output the final merged table as a downloadable CSV or as a formatted table you can paste into your calibration deck.
What you should see: A CSV file in your downloads folder, still keyed to ID codes, ready to be matched back to real names in a secure system.
Building the merged tier view:
Merge these two files on the ID column and group into [your tier names]. Show ID,
rating, and comp position for each row, sorted by tier.
Finding comp-to-rating mismatches:
Flag any employee with a rating in the top tier but a comp band percentile below
[your threshold]. List ID and the specific numbers that triggered the flag.
Checking for stalled ratings:
If the file includes rating history, flag any employee with the same rating for
three or more consecutive review cycles and no comp change in that window.
Getting a plain-language summary for the calibration lead:
Summarize this tier breakdown in three sentences a calibration facilitator could
read aloud to open the meeting, without naming any individual employee.
Preparing a devil's-advocate check:
Look at the flagged list again and suggest one alternative, performance-based
explanation for each flag that isn't about comp or tenure, so I can raise it as
a discussion point rather than a conclusion.