One PDF, hashed forever.
The moment a file lands, we compute a SHA-256 of its contents. If that exact hash already exists in our knowledge graph, we skip parsing entirely and link to the prior extraction — saving the whole OCR + topic-extraction pipeline.