Where is your candidate data stored?
Your CRM vendor can answer for two of the places it lives. The rest are copies your own agency made, and a client questionnaire will eventually ask about all of them.
When a client sends a data questionnaire, the usual reflex is to forward the hard parts to the CRM vendor and wait. That gets you an accurate answer to a smaller question than the one being asked, because a candidate's details do not sit in one place. They sit in several, most of which your agency created itself, and the ones outside the CRM are where almost every real problem starts.
This is the map I would draw for any agency owner who wants to answer that questionnaire honestly rather than quickly.
What does a CRM vendor actually mean by storage?
A vendor answering that question is usually describing two things: the database holding your records, and the object storage holding your files. The database has the structured detail, meaning names, contact details, work history, notes, statuses and everything typed into a field. The object storage has the CVs, right-to-work documents, recordings and anything else uploaded as a file.
Three further places exist in almost every product and are less often volunteered. Backups, which are separate copies held to a schedule and kept for a defined period. Search indexes, which hold a processed copy of your data so that searching is fast. And logs, which record what happened and frequently contain identifiers, sometimes more.
Ask about all five by name. A vendor who can describe each one, including how long backups are kept and what is in the logs, has thought about this properly. A vendor who answers only about the database is not hiding anything, they are answering the question you asked.
Which region is it in, and what can move it?
Region is decided when your account is created and it is one of the few things that is genuinely difficult to change later, so it belongs in the conversation before you sign rather than in the questionnaire afterwards. If any of your clients have a requirement about where candidate data sits, get the answer in writing at that point.
The subtlety worth knowing is that different parts of a product can live in different places. The database may sit in one region while a support system, an analytics tool or a model provider sits elsewhere. That is normal and it is not automatically a problem. It becomes a problem when nobody can list which parts go where, since that is precisely what a serious client will ask.
There is a second question behind the region, which is who else can reach it. Every hosting arrangement involves a provider, and the honest position is that your vendor holds data on infrastructure somebody else operates. What you want to know is the list of suppliers who process your data and what each of them does with it, and any vendor selling to agencies in the UK, Europe or the Gulf should be able to hand that list over without a project starting.
The copies your own agency made
Most candidate data outside your CRM was put there by your own team doing ordinary work, and it is the part no supplier can account for. Email is the largest of these. Every CV sent to a client sits in two mailboxes indefinitely, along with everything attached to every thread, and mail is rarely covered by any retention policy an agency has written down.
Then there are the accounts you hold elsewhere. Job boards keep the applicants who applied through them. Sourcing tools keep the lists and projects your consultants built. Assessment and checking suppliers keep their own records of every candidate you put through them. Each of those is a separate holder of the same person's details with its own retention behaviour.
The two that cause the most trouble are the smallest. Exports to a spreadsheet, made for a perfectly good reason and then left in a shared drive or a downloads folder forever. And candidate details on consultants' phones, in personal messaging apps and in saved contacts, which is how a leaver takes a desk with them and is far more common than anything an attacker does.
What happens when AI reads a CV
When AI processes a CV, that text is sent to a model provider, processed and returned, and the questions that matter are which provider, whether anything is retained there, and whether your data is used for training. Those three questions have short answers at any vendor that has thought about it, and vagueness on the training question in particular is worth pressing.
Ask a fourth question that people forget. Ask whether the vendor can list every place in their product where a model sees candidate data. A list kept by hand goes stale immediately, so what you are really asking is whether they have a way of knowing, and the answer separates the products where this was designed from the ones where AI was added feature by feature.
The related practical question is what the AI writes back. Ours reports how sure it is on every judgement and hands anything below three-quarters confidence to a recruiter instead of writing it to the record, which matters here because an AI that guesses quietly is also an AI that puts unverified personal data onto a candidate's file. Accuracy and data protection turn out to be the same conversation.
| Where a copy lives | Who can answer | Ask |
|---|---|---|
| Records database | Vendor | Which region, and can it change |
| Files, CVs, recordings | Vendor | Same region, and who can download them |
| Backups | Vendor | How long kept, and does deletion reach them |
| Search index and logs | Vendor | What personal data is in them, kept how long |
| Model providers | Vendor | Who, retained where, used for training |
| Mailboxes | You | What is your retention rule for sent CVs |
| Job boards and sourcing tools | You | What do they keep after you cancel |
| Exports and spreadsheets | You | Who can export, and where do files land |
| Phones and chat apps | You | What leaves with a consultant who resigns |
What happens when you delete a candidate
Deleting a record in a CRM removes it from your view immediately and removes it from every copy over a longer period, and the difference is worth understanding before you promise anything to a candidate. Backups are the clearest example: a copy taken last week still contains the person until that backup expires, which is normal practice at every vendor and is the reason retention periods are stated rather than infinite.
Ask three specific things. Whether deletion is a genuine removal or a hidden flag on a row that still exists. How long until it reaches backups and indexes. And whether audit entries about the record survive, which they often must, so that there is still evidence of who did what.
Then do the harder half yourself. A deletion request from a candidate has to reach your mailboxes, your job board accounts, your sourcing tools and any spreadsheet somebody made, and none of that happens automatically. Agencies that handle this well have a written list of every place candidate data goes, kept by a named person, and they check it once a year.
The five-minute version for a questionnaire
Keep one page listing every system that holds candidate data, who supplies it, which region it sits in, how long it keeps things, and who at your agency can export from it. Build it once and most questionnaires become a copy-and-paste job instead of a fortnight of chasing.
My own opinion is that agencies worry about the wrong end of this. The vendor's storage arrangements are the part most likely to be competently handled, because it is a supplier's core job and a client will audit it. The uncontrolled copies, meaning the exports, the mailboxes and the phones, are the part nobody owns, and every serious candidate data problem I have watched an agency deal with started in that column rather than in a database.
Siva is an engineer at Recruitly.



