Secludy AI — Make sensitive data safe to use for AI

Make sensitive data safe to use for AI

GraphReplica replaces the sensitive PII entities in your data with realistic synthetic data.

The same person or entity gets the same replacement across your tables, documents, and images in one pass. One replacement, matched across all data formats.

Format Description
Tables Data structured in rows and columns
Documents Main written content and text containers
Images Visual representations or graphics

Original

Replica

Zero leakage you can prove

Train a model on raw enterprise data and it memorizes what it sees. In stress tests, unprotected models leaked about 27.5% of injected sensitive values. Data built with GraphReplica leaked none. Every run ships an audit-ready report you can hand to your legal and security teams.

PII leakage stress test

Model Leakage
Unprotected model 27.5%
GraphReplica 0%

Use cases

Put safe data to work

GraphReplica unblocks the work that real data used to block. One safe replica that stays realistic and usable.

Train and evaluate AI agents

Build realistic environments to train, evaluate and red-team agents on data that behaves like production.

Unblock coding agents and BI

Point coding tools and BI at a safe replica instead of waiting on privacy review.

Safe demos, QA and staging

Stand up demo, test and staging data without exposing real customers.

License and sell data

Sell or license datasets to AI labs and partners without exposing PII, PHI or IP.

Keep joins across tables

Keep foreign keys and relationships intact across many tables and files.

Resolve identity across docs

Match the same real entity across email, PDFs, spreadsheets and notes.

Test for PII and PHI leakage

Prove whether sensitive data leaked with membership inference and canary tests.

Move past privacy review

Ship AI work in days instead of waiting on long privacy reviews.

Interactive demo

Try GraphReplica

Step through a graph-preserved replica, then move the slider to see why masking breaks data and replicas keep it usable.

Input formats

Format Example
XLSX Excel / CSV
PDF Resume
JSONL Recruiter notes { candidate: "Sofia Martinez", zip: "94107" }
TXT Interview feedback

Privacy-aware entity graph

Original Replica
Sofia Martinez pending
sofia.martinez@example.edu pending
94107 pending
Bayview University pending
Northstar Robotics pending
ML Infrastructure Intern pending

Why teams choose GraphReplica

Same entity, same stand-in, everywhere

GraphReplica finds the sensitive entities in your data and replaces only those. The same real person, customer, employee or account gets the same stand-in across every file, table and document. This holds across millions of records and years of history. Random replacement breaks this. Masking leaves nothing usable.

Joins survive

Foreign keys and relationships stay intact across many tables and documents. Your downstream joins and queries still work.

Runs in your environment

A container that runs in your cloud, data center or Databricks. Your data never leaves. Every run is air-gapped and audit-ready.

Unstructured data

Works on free text too. Toggle between the source and the safe replica.

Compliance

Built to meet data protection requirements across the EU, US and APAC with one integration. Built to meet GDPR, CCPA and HIPAA requirements. No customer data is retained after processing.

How it works

From messy data to a safe replica

GraphReplica runs five stages and gates every release on a leak check. No release ships if an original value survives.

  1. Detect
    Find sensitive entities across messy multi-format sources.
  2. Resolve
    Group the records that refer to the same real entity. Surface conflicts.
  3. Replace
    Swap only the sensitive entities for consistent realistic stand-ins.
  4. Validate
    Run a leak check and a consistency check on the output.
  5. Report
    Produce audit-ready detection, replacement and risk reports.

0% PII leaked in stress tests
100M+ records held consistent
0.9 F1 detection and replacement
Within 5% of real-data utility

Support

Frequently asked questions

What does GraphReplica do?

GraphReplica finds the sensitive entities in your data and replaces only those with realistic stand-ins. Everything that is not sensitive stays exactly as it was. The same real entity gets the same stand-in across every file, table and document.

How is this different from data masking?

Masking removes values and leaves your data unusable. GraphReplica swaps sensitive values for realistic stand-ins so the data still reads naturally and your downstream work still runs.

Does my data ever leave my environment?

No. GraphReplica runs as a container in your cloud, data center or Databricks. Your data never reaches Secludy. Every run is air-gapped.

Which regulations does this help with?

GraphReplica is built to meet GDPR, CCPA and HIPAA requirements. No customer data is retained after processing.