Solved by Remove PDF metadata
This feature helps prevent search engines and automated scrapers from collecting embedded PDF metadata fields when you publish PDFs on your website. It reduces unintended exposure of author names, software details, and other hidden document properties that can be harvested at scale.
When you upload PDFs to a website, the document may include embedded metadata such as author, organization, document title, creation tools, timestamps, keywords, and custom fields. This feature is designed to reduce the chance that search engines and scraping tools can collect those metadata fields from your published PDFs. It focuses on minimizing what can be harvested from the PDF file’s internal properties, helping you avoid inadvertent disclosure of internal names, systems, or workflow details. Use it when preparing PDFs for public access so that only the visible content is shared, not the hidden document properties. This is especially useful for public reports, downloadable brochures, policy documents, or any file that might be mirrored, cached, or bulk-collected. The feature supports privacy and security best practices by limiting passive information leakage that can be used for profiling or reconnaissance. It can also help with compliance requirements where personal data or internal identifiers should not be distributed. For teams, it provides a consistent way to sanitize documents before publishing, reducing manual effort and the risk of missing a field. The result is cleaner, safer PDF publishing without changing the visible content of the documents.
External Resource
https://cross-service-solutions.com/
If you know of a tool or approach that could help people solve a problem we haven't covered yet, we'd love to hear about it.