Content-hash deduplication identifies duplicate records by hashing their content — typically sha256 — instead of comparing titles, filenames, or metadata. Two records are the same asset only if the bytes agree.
Title matching finds candidates, not duplicates. In a live audit, 24 of 30 title-matched "duplicate" groups dissolved under content hashing: same-title different conversations, test probes differing only by an embedded marker, and rows whose provenance field collided while their content was unrelated.
Full-content hash first; stripped-body similarity second; provenance third; creation-time gap fourth. And when a row has no content at all, it cannot be classified — quarantine, not merge.
This entry is part of the untactit glossary. Definitions draw on measured write-ups published on the blog.
Connect one workspace and see every skill, rule, and memory your team has in play — in about ten minutes.
No credit card. Works with what you already run.