What Confidential Documents Can Google Find About Your Business?

Google can index publicly accessible business documents. Organisations should search authorised domains, secure the original files, assess personal-data exposure and monitor recurrence. Removing a result from search does not secure the original file.
A business has operated several websites over many years. Staff have uploaded brochures, quotations, tender documents, meeting packs, customer forms, spreadsheets, technical diagrams, staff handbooks, presentations, archived reports and scanned letters. Some were intended to be public. Others were uploaded temporarily and forgotten.
Those documents may now sit on the current website, an old website, a test site, a former supplier's server, a public SharePoint or OneDrive link, a Google Drive sharing link, a marketing platform, an old subdomain or a downloadable-media folder.
A simple authorised search may reveal material nobody currently remembers publishing — and may reveal it to anyone with internet access.
The document exposure usually happened when the file was published — not when somebody later found it through Google.
Google does not need to bypass security to index a document that the website already serves publicly. If a file can be opened by anyone without signing in, it is already publicly accessible. The search engine is simply making it easier to find.
If a document can be opened without signing in, you should assume it may eventually be discovered, indexed or shared.
The Quick Answer
How to Check Your Domain for Indexed Documents
Businesses can use Google's site: and filetype: search operators to look for documents indexed under domains they own or are authorised to assess.
Examples:
- site:example.co.uk filetype:pdf — finds indexed PDF documents
- site:example.co.uk filetype:xlsx — finds indexed Excel workbooks
- site:example.co.uk filetype:docx — finds indexed Word documents
- site:example.co.uk filetype:pptx — finds indexed PowerPoint presentations
- site:example.co.uk "confidential" — searches indexed content for that phrase
These searches may identify legitimate brochures and reports, old documents, duplicate files, internal-looking material, files containing personal data and documents published from obsolete systems.
However, search results are incomplete, a missing result does not prove that no exposure exists, documents may remain directly accessible even after search removal, and every finding requires human review.
Only search domains and systems you own or are authorised to assess. Use search as a discovery check — not as proof that the organisation has no exposed documents.
How Search Engines Find Business Documents
A search engine may discover a business document without any deliberate act by the organisation. Discovery can occur through a link from the website itself, a link from another website, a public sitemap, a public directory listing, an old indexed page, a social-media post, an email or newsletter archive, a public cloud-sharing link, URL pattern discovery or an embedded reference within another document.
The document does not need to appear in the main navigation for it to be indexed. It does not need to be prominently linked. It does not even need to be intentionally published. If the file is publicly accessible at a URL the crawler discovers, it may be indexed.
A file can be absent from the menu and still be fully public.
Hidden from navigation is not the same as protected from access.
Which File Types Google Can Index
Google currently says it can index the contents of many file formats beyond standard web pages. Currently documented types include PDF, Microsoft Word (DOC and DOCX), Microsoft Excel (XLS and XLSX), Microsoft PowerPoint (PPT and PPTX), Rich Text Format (RTF), OpenDocument formats (ODT, ODS, ODP), EPUB and PostScript. Verify the complete current list at Google Search Central before relying on this information.
Last checked: 2 August 2026. Verify the current list of indexable file types at developers.google.com/search/docs.
When Google indexes a document, it can expose file titles, visible body text, table contents, document metadata, snippet extracts, names, figures and contact details. Not every hidden element is guaranteed to be extracted or displayed in search results, but the visible text of an indexed document may appear in full in the search index.
Uploading a document to a public website can make its contents searchable — not merely the filename.
Safe Searches to Run Against Your Own Domain
The following searches are for use against domains and systems the organisation owns or has explicit authority to assess. Do not use these searches against unrelated organisations' domains, and do not attempt to access, download or circulate any document you find beyond what is necessary for the assessment.
| Search purpose | Search query |
|---|---|
| All indexed results on the domain | site:example.co.uk |
| All indexed PDF documents | site:example.co.uk filetype:pdf |
| Microsoft Word documents (.docx) | site:example.co.uk filetype:docx |
| Older Microsoft Word files (.doc) | site:example.co.uk filetype:doc |
| Microsoft Excel workbooks (.xlsx) | site:example.co.uk filetype:xlsx |
| Older Microsoft Excel files (.xls) | site:example.co.uk filetype:xls |
| Microsoft PowerPoint presentations (.pptx) | site:example.co.uk filetype:pptx |
| Indexed content containing "confidential" | site:example.co.uk "confidential" |
| Indexed content containing "internal use" | site:example.co.uk "internal use" |
| Indexed content containing staff-type terms | site:example.co.uk "employee" |
Replace example.co.uk with the actual domain being assessed. Also check authorised previous business domains, historic trading names, test subdomains, portal subdomains, media or uploads subdomains, supplier-hosted domains and public cloud-sharing domains.
For more reliable administrative information on domains the organisation controls, use Google Search Console tools including URL Inspection, the Page Indexing report and the Removals tool. These provide data that public searches alone cannot.
Do not attempt to authenticate to, bypass access controls on or access files from domains and systems you do not own or have explicit written authority to assess.
What the Results May Reveal
Not every indexed document is a problem. A search may return a mix of legitimate public content and material that warrants further review.
| Likely legitimate | Warrants closer review |
|---|---|
| Published brochures and marketing materials | Customer lists or customer contact details |
| Annual reports and public accounts | Employee names, addresses or pay details |
| Public policies (privacy notice, cookie policy) | Internal pricing, quotes or tender responses |
| Published price lists for public use | Signed or unsigned contracts |
| Product or service instructions | Supplier terms and commercial agreements |
| Public-tender submissions by design | Network diagrams or technical configuration |
| Regulatory submissions | Meeting minutes or internal project plans |
| Press releases and news items | Scanned identification or payroll documents |
| Bank details or credentials | |
| Security assessments or vulnerability reports | |
| HR, disciplinary or personal reference documents |
Context determines risk. A document titled 'confidential' does not itself prove a breach. A staff handbook containing employee names may be entirely legitimate if it was intended to be public. A public policy containing a named data-protection officer is not an exposure.
The purpose of the review is classification and judgement — not counting every indexed file as an incident.
Search Limitations
What Google search results may miss
Google's search operators are useful discovery tools. They are not comprehensive audits. Files may be absent from search results because they have not yet been crawled, are indexed under another URL, the result has been omitted for other reasons, the content was recently published, another search engine indexed it, the file is only reachable through a known direct link, or the content sits on a supplier's domain.
Search results may also include duplicate files, outdated snippets, removed pages awaiting recrawling and public documents that are entirely legitimate.
No search result is not evidence of no exposure.
For a thorough review, combine public search with Google Search Console, website file inventories, hosting reviews, cloud-sharing reports, Microsoft 365 sharing reviews, Google Workspace sharing reviews, DNS and subdomain inventories, former-supplier checks, staff interviews and backup and archive reviews.
Public Website Files Versus Secure Cloud Documents
Cloud storage does not automatically make a file private. Privacy depends entirely on the link and permission configuration, not on where the file is stored.
| Access type | What it means |
|---|---|
| Public website file | Hosted openly and accessible without signing in. Can be indexed and found by anyone. |
| Public share link | A cloud-hosted file configured so anyone with the link, or anyone on the internet, can access it. May be indexed if discovered. |
| Organisation-only link | Requires a valid account from the organisation before access is granted. Not publicly accessible. |
| Named-user link | Only specified authorised individuals can open it. The most restrictive setting. |
SharePoint, OneDrive, Teams, Google Drive, Dropbox and document portals all offer public sharing options. A file shared with 'Anyone with the link' on any of these platforms is effectively public. If that link is discovered, emailed, posted online or included in a sitemap, it may be indexed.
'Stored in the cloud' describes location. It does not describe access control.
Common Causes of Document Exposure
Most document exposures are accidental rather than the result of a deliberate decision. Common causes include:
- Uploading an internal file to the website's media library rather than an internal system
- Publishing the wrong attachment — attaching an internal version instead of a public one
- Leaving an old document online when it was no longer current
- Cloud sharing set to 'anyone with the link' when the intention was named access
- Sending a public cloud link rather than a named-user link
- Copying files from a development site that retained internal documents
- Retaining an obsolete website domain without an access review
- Forgotten test subdomains remaining live and publicly accessible
- Former web agencies retaining files on servers no longer under the organisation's control
- Insecure document portals with inadequate authentication
- Publishing meeting packs as PDF files on a publicly accessible subdomain
- Public backup folders or export directories accessible from the web
- Exports placed in web-accessible directories by automated processes
- Weak content-management permissions allowing staff to publish without review
- Assuming an obscure or complex URL is effectively private
- Failed redaction — covering text visually without removing the underlying data
- Hidden spreadsheet tabs or rows that remain in the published file
- Tracked changes or comments not removed before publication
- Document metadata including author names, previous filenames and revision history
- Filenames containing customer or employee personal information
Metadata, Comments and Hidden Information
A document may contain significantly more information than the visible page reveals. Someone downloading a publicly accessible file may inspect more than the search snippet shows.
| Source | What may be exposed |
|---|---|
| Document properties | Author name, company name, previous filenames, revision count, creation date |
| Tracked changes | Previous drafts, deleted text, names of reviewers and dates of changes |
| Comments | Internal reviewer notes, names, contact details, decision rationale |
| Hidden worksheets | Data excluded from the visible sheet but still present in the file |
| Hidden rows and columns | Filtered-out data still present in the workbook |
| Speaker notes | Presentation speaker notes not visible on printed slides |
| Embedded attachments | Other files attached inside a document container |
| Image metadata | Camera or device metadata embedded in images within the document |
| Formulas and linked data | Source data, server paths or external references visible in formula cells |
| Visual redaction errors | Text that appears covered but remains selectable and copyable as underlying data |
A search result may reveal the document; downloading it may reveal information the publisher never intended to include.
Redaction: Visual Covering Versus True Removal
Placing a black rectangle or colour block over text in a document does not remove the underlying information. A recipient can often select, copy and paste the covered text from a PDF or copy the covered cell value from a spreadsheet. True redaction requires permanently removing the underlying content from the published copy.
Recommended redaction practice: use an approved redaction tool that permanently removes the selected content, create a separate public copy rather than modifying the original, flatten or sanitise the document where appropriate, strip metadata, inspect the final exported file directly, test by attempting to copy text from the redacted region, check hidden sheets and comments, have another person review the sanitised copy and retain the controlled original in a secure, non-public location.
What to Do When You Find a Sensitive File
If an authorised review reveals a document that should not have been publicly accessible, follow a structured response sequence.
- 1. Do not circulate it unnecessarily — do not forward, share or copy the document beyond what is needed to manage the incident
- 2. Record the finding — capture the URL, the search query used, the date and time, the file type, the responsible system and the observed access conditions. Do not reproduce excessive personal data in the incident record
- 3. Restrict or remove the source — depending on the situation, delete the file, require authentication, correct sharing permissions, replace the file with a sanitised version or configure the server to return an appropriate unavailable response
- 4. Check for copies — review duplicate URLs, old domains, cloud links, caches, mirrors, supplier systems and archive copies
- 5. Use search-removal tools — where appropriate, use Google Search Console to request temporary removal or refreshed indexing of the URL
- 6. Assess the information — identify what personal data, commercial information, credentials or customer data was involved
- 7. Contain related risk — revoke exposed credentials, change passwords, replace API keys, notify affected suppliers and monitor for misuse
- 8. Escalate — inform the data-protection lead, business owner, IT provider, legal adviser, insurer and regulator where required
- 9. Verify — confirm that direct access is blocked, sharing permissions are corrected, search results update appropriately and duplicate files have been addressed
- 10. Document and review — record lessons learned and identify preventive actions for future documents
Remove the source first. Search-result removal is an additional containment step, not the permanent fix.
Removing a Search Result Versus Securing the Source
| Method | What it does | What it does not do |
|---|---|---|
| Google Removals tool | May temporarily suppress eligible URLs from Google Search results for a defined period | Delete the source file, remove it from other search engines, stop direct access, invalidate copied links or revoke previous downloads |
| noindex | Instructs supporting search engines not to include content in their index after the next crawl | Require authentication, stop direct access, prevent someone sharing the URL or guarantee compliance by search engines that disregard the instruction |
| robots.txt | Controls crawler access — may prevent some crawlers from visiting the URL | Secure access to the file, prevent indexing where a URL is discovered through other means or constitute a confidentiality mechanism of any kind |
| Password or identity-based access | Restricts access to authenticated and authorised users | Remove existing indexed snippets or previously downloaded copies |
| Source deletion | Removes the file from the host where it no longer needs to exist publicly | Remove cached or archived copies held elsewhere |
Search controls manage discoverability. Authentication manages access.
Personal Data and Incident Response
Publicly exposing personal information without appropriate authorisation may constitute a personal-data breach under UK data-protection law. A security breach causing accidental or unlawful disclosure of, or access to, personal data is a breach that the organisation is required to assess and potentially report.
The organisation should assess what personal data was involved, how many individuals were affected, how long the file was accessible, whether it was indexed, whether access logs exist, whether downloads occurred, the sensitivity of the information, the possible harm to individuals, any contractual obligations and relevant notification requirements.
UK organisations may have time-sensitive obligations to report to the ICO. The ICO currently requires reporting of personal-data breaches likely to result in a risk to individuals' rights and freedoms within 72 hours of becoming aware. Whether notification is required depends on the likelihood and severity of risk to individuals. Obtain appropriate data-protection or legal advice rather than assuming that removal from search alone closes the incident.
Last checked: 2 August 2026. Verify current ICO breach-reporting requirements at ico.org.uk.
Preventing Future Exposure
Most document exposures are preventable through clear process, training and periodic review. Recommended controls include:
- Clear information classification — define what is public, internal, confidential and restricted before publication
- Approved publishing processes — require review and sign-off before any file is published publicly
- Separate public and internal document libraries — do not upload internal files to website media folders
- Named ownership of website uploads — every published file should have an identifiable owner
- Review before publication — check content, metadata, hidden elements and redaction before uploading
- True redaction — use approved tools that permanently remove underlying content
- Metadata removal — strip author, revision and embedded data from files before public publication
- Sharing-link controls — use named-user links rather than public links wherever possible
- Expiry dates for public links where the platform supports them
- Periodic access reviews — regularly check which documents are publicly accessible
- Website file inventories — maintain a record of what has been published
- Old-domain retirement — take down or password-protect domains no longer in active use
- Subdomain inventories — record all active subdomains and their purpose
- Google Search Console monitoring — use the Page Indexing and Coverage reports regularly
- Microsoft 365 sharing reports — use the SharePoint admin centre and Purview to identify overly-permissive sharing
- Google Workspace sharing reviews — audit Drive sharing settings and external link counts
- Staff training — ensure staff understand document classification and publishing permissions
- Incident reporting — provide a clear route for staff to report accidental publication
- Supplier exit checks — require former web providers to confirm and document file deletion
The safest publishing workflow assumes every public document may be indexed, downloaded and retained.
Running a Document Exposure Review
Six-Phase Review Process
| Phase | Actions |
|---|---|
| Phase 1 — Identify | Record current domains, previous domains, subdomains, website hosts, cloud platforms, document portals, marketing systems and external supplier-hosted systems |
| Phase 2 — Search | Run authorised site and filetype searches, exact business phrases, Search Console indexing reports, website file inventories and cloud-sharing reports |
| Phase 3 — Classify | Classify each finding as: intended public, outdated public, internal, confidential, restricted or unknown |
| Phase 4 — Remediate | For each finding, decide to retain, update, replace with sanitised version, restrict access, delete or escalate to a senior decision-maker |
| Phase 5 — Verify | Confirm that direct access is blocked, sharing permissions are corrected, search status is updated, duplicate copies have been addressed and decisions are documented |
| Phase 6 — Maintain | Set a recurring review date and assign ownership of each domain and system reviewed |
What Can the Internet Find About Your Business?
Document Exposure Checklist
- □ Have we searched our current domain with site:?
- □ Have we searched old and historic domains?
- □ Have we searched for PDF documents?
- □ Have we searched for Word documents?
- □ Have we searched for spreadsheets?
- □ Have we searched for presentations?
- □ Have we checked test sites and subdomains?
- □ Have we reviewed public SharePoint links?
- □ Have we reviewed public OneDrive links?
- □ Have we reviewed Google Drive sharing settings?
- □ Have we checked website upload and media folders?
- □ Have we checked former web providers' systems?
- □ Do we know who approves files before publication?
- □ Are documents properly redacted (not just visually covered)?
- □ Are metadata and comments removed before publication?
- □ Are hidden spreadsheet tabs and rows checked?
- □ Do public links have expiry dates where possible?
- □ Do we have a documented data-breach response process?
- □ Does each online system have a named owner?
- □ When will the next review occur?
Warning Signs
The organisation may have an exposure problem where:
- Search reveals internal-looking documents publicly accessible
- Public documents contain personal data that was not intended for publication
- Filenames include customer or employee names
- The same file exists under several different URLs
- Old domains remain active and contain unreviewed content
- Test or development sites appear in search results
- Cloud sharing uses 'anyone with the link' without review
- Nobody owns or monitors the website media library
- Documents have been visually covered rather than genuinely redacted
- Hidden spreadsheet tabs remain in published files
- Presentations contain unreviewed speaker notes
- Former suppliers still host files on servers no longer under the organisation's control
- Google removal is treated as the complete and only fix
- robots.txt is being used as a substitute for proper access control
- Confidential files rely on an obscure URL for protection
- Credentials or API keys appear in downloadable documents
- No incident process exists to handle an accidental publication
An obscure link is not an access-control system.
Practical Business Implications
| Implication | What it means in practice |
|---|---|
| Search engines can index business documents | Public files may become searchable by content and file type, not merely by filename. |
| Old files create long-term risk | A document can remain online and indexed long after its original purpose ended. |
| Public cloud links need review | Cloud hosting does not guarantee restricted access. Permissions must be checked explicitly. |
| Search results are only part of the picture | Direct links, other search engines and previously downloaded copies may remain regardless of Google removal. |
| Removal must address the source | Suppressing a Google result is not permanent protection. The file must be secured at source. |
| Personal data may create legal duties | Accidental exposure of personal information may require formal breach assessment and possibly ICO notification. |
| Document hygiene matters | Metadata, comments, hidden sheets and visual-only redaction can expose additional information beyond the visible content. |
| Regular reviews are necessary | Websites, suppliers and sharing permissions change over time. A one-time check is not sufficient. |
The IT Club View
This is one of the simplest cyber-security checks a business can perform. It requires no specialist scanning tool to begin, no external access and no technical background beyond understanding how to enter a search query.
A few authorised searches may reveal forgotten files, obsolete documents, public cloud links that were never reviewed, poor or failed redaction, old websites nobody remembered were still live and documents with unclear or absent ownership. That is valuable and actionable intelligence for any organisation.
But the search itself is only discovery. Finding a document through Google does not mean Google created the problem. The organisation still needs to classify the finding, remove or restrict the source, assess any personal-data breach, correct the process that allowed the exposure, check for duplicate copies and monitor for recurrence.
Google is often the messenger. The underlying problem is that the document was accessible without appropriate control.
IT Club recommends checking owned domains regularly, reviewing old domains and subdomains, inspecting public cloud-sharing links, separating public and internal document storage, treating robots.txt as crawler control rather than security, using proper authentication for anything that should not be public, applying genuine redaction, maintaining a documented incident-response process, assigning named ownership for every domain and document library and repeating the check on a scheduled basis through an Operational Heartbeat.
Search for your own information before somebody with less helpful intentions does — but fix the access, not merely the search result.
Plain-English Takeaway
Google can index publicly accessible PDFs, Word documents, spreadsheets, presentations and other files. Businesses can use authorised site and file-type searches to identify documents associated with their own domains, but search results are incomplete and removal from Google does not secure the original file. Sensitive content should be deleted, corrected or placed behind proper authentication, with any personal-data exposure assessed through the organisation's incident-response process.
Related Business Questions
Can Google index PDF files?
Yes. Google currently says it can index the content of publicly accessible PDF files, including their visible text, tables and some metadata. A PDF uploaded to a public website may appear in search results and have its content searchable.
Can Google read Word documents?
Google currently says it can index the content of Microsoft Word files (both DOC and DOCX format) that are publicly accessible. Text, tables and visible content may appear in the search index.
Can Google index Excel spreadsheets?
Google currently says it can index the content of Microsoft Excel files (XLS and XLSX format) that are publicly accessible. Visible cell values, sheet names and document properties may appear in the index.
Can Google index PowerPoint presentations?
Google currently says it can index the content of Microsoft PowerPoint files (PPT and PPTX format) that are publicly accessible. Slide text, speaker notes and visible content may appear in the index.
How do I search my website for PDFs?
Use the search query: site:yourdomain.co.uk filetype:pdf — replacing yourdomain.co.uk with your actual domain. This asks Google to return any PDF files it has indexed under that domain. Only use this on domains you own or are authorised to assess.
What does the site: search operator do?
The site: operator restricts Google search results to a specified domain or subdomain. For example, site:example.co.uk returns only results Google has indexed under that domain. It does not return every file on the domain — only those Google has indexed.
What does the filetype: operator do?
The filetype: operator restricts search results to a specific file extension. Combined with site:, it allows a search such as site:example.co.uk filetype:pdf which returns only PDF files Google has indexed under that domain.
Are Google search operators complete?
No. Google states that search operators are constrained by indexing and retrieval limits. A clean result set does not mean no files exist. Files may be accessible but not yet crawled, indexed under different URLs, excluded from results for other reasons or reachable only via a known direct link.
Does no search result mean a file is private?
No. A file absent from Google search results may still be directly accessible to anyone with the URL. It may also be indexed by other search engines or accessible via archived copies. No search result is not evidence of no exposure.
Can an unlinked file appear in Google?
Potentially yes. If the URL is discovered through a sitemap, another website, a social-media post, an email archive or URL pattern recognition, Google may crawl and index it even without a prominent link from the main website.
Can Google index OneDrive links?
A OneDrive file configured with 'Anyone with the link' sharing or public internet access may be indexed if the URL is discovered by a search engine. OneDrive files configured for organisation-only or named-user access should not be publicly accessible.
Can Google index SharePoint documents?
SharePoint files configured for anonymous or public internet access may be indexed. Files within SharePoint requiring organisational authentication should not be publicly accessible, but incorrectly configured sharing settings have been a common source of exposure.
Can Google index Google Drive files?
Google Drive files shared with 'Anyone with the link' or published to the web may be indexed. Files restricted to specific users or the organisation should not be publicly accessible.
Is 'anyone with the link' private?
No. 'Anyone with the link' sharing means any person who obtains the URL can access the file without signing in. If the URL is shared, forwarded, posted or discovered, the file is accessible to anyone who has it.
Does robots.txt protect confidential documents?
No. Google says robots.txt is a crawler-management mechanism. It instructs supporting crawlers not to visit certain URLs. A blocked URL may still be indexed if Google discovers the address from another source. robots.txt does not restrict direct access to the file by a person.
What does noindex do?
The noindex directive instructs supporting search engines not to include the item in their search index. It may reduce the chance of the document appearing in search results once processed.
Does noindex prevent direct access?
No. noindex affects search indexing. It does not restrict who can open the file. Anyone who knows or discovers the URL can still access a noindex-tagged document if it is publicly hosted.
How do I remove a document from Google?
For domains you own, use Google Search Console's Removals tool to request temporary suppression of a URL from Google Search. This requires verified ownership in Search Console. The suppression is temporary — securing or removing the source file is required for a permanent fix.
Does Google removal delete the original file?
No. The Google Removals tool temporarily suppresses the URL from appearing in Google Search results. It does not delete the source file, remove it from the web server, stop direct access or affect other search engines.
How long does a temporary removal last?
Google currently says temporary removals suppress the URL for approximately six months. After that period, the URL may reappear in results if the file is still publicly accessible. Verify the current duration at Google Search Central.
How do I permanently remove a document?
Permanent removal requires action at the source: delete the file, require authentication to access it, configure the server to return a 404 or 410 status, apply a noindex tag if appropriate and request a recrawl via Search Console. Once Google confirms the page returns an unavailable response, the removal becomes permanent.
Can a deleted file still appear in search?
Yes, temporarily. After a file is deleted, Google may continue to show a cached or indexed result until it recrawls the URL and confirms the file is gone. Use the Search Console Removals tool to accelerate suppression while the index updates.
What is the outdated-content tool?
Google Search Central provides a tool for requesting removal of cached or outdated content where the content on the live page has changed but cached results still show old information. It is not for the same purpose as the Removals tool, which addresses current indexed content.
What is true document redaction?
True redaction permanently removes the underlying content from the published copy of a document. An approved redaction tool deletes the selected text or image from the file data so it cannot be retrieved. The result is typically a flattened, sanitised copy suitable for public release.
Is drawing a black box over text safe?
No. Placing a black shape, highlight or filled rectangle over text in a PDF or Word document typically does not remove the underlying data. A recipient can often select the text beneath, copy it or use accessibility tools to retrieve it.
Can hidden Excel sheets be exposed?
Yes. An Excel workbook published publicly retains all hidden sheets unless they are removed or the file is exported to a specific format that excludes them. A recipient can unhide sheets in the downloaded file. The same applies to hidden rows and columns.
Can Word comments remain in a published document?
Yes. Word documents retain tracked changes, comments and revision history unless explicitly accepted, deleted and removed before saving. Accepting all changes and deleting all comments before exporting the public copy is recommended.
Can PowerPoint notes be exposed?
Speaker notes are stored in a PowerPoint file and may be accessible to anyone who downloads it. Notes panels are not part of the visible presentation but are present in the file data. Remove or review speaker notes before publishing a PowerPoint file publicly.
What document metadata should be removed?
Before publicly publishing a document, consider removing author name, company name, previous filenames, revision history, creation and edit dates, comments and tracked changes. Most Office applications provide an Inspect Document function to identify and remove hidden data before publication.
What should I do if personal data is indexed?
Restrict or remove the source file, request suppression via Google Search Console, assess what personal data was involved and for how long, evaluate the risk to individuals affected, document the incident and obtain data-protection advice to determine whether ICO notification or individual notification is required.
Is an indexed document a data breach?
An indexed document containing personal data may constitute a personal-data breach. A breach is a security incident causing accidental or unlawful disclosure of or access to personal data. Whether it is reportable depends on the likelihood and severity of risk to individuals — not every exposure automatically requires ICO notification.
When must a breach be reported to the ICO?
Under UK GDPR, reportable personal-data breaches must be reported to the ICO within 72 hours of the organisation becoming aware, where the breach is likely to result in a risk to individuals' rights and freedoms. Obtain data-protection advice rather than making assumptions about reportability. Verify current ICO guidance at ico.org.uk.
How should businesses review old websites?
Compile an inventory of all past domains and subdomains. Run site: searches against each. Check with current and former web providers for file inventories. Review whether authentication or a holding page has been put in place. Verify who currently controls the hosting and DNS for each domain.
Can former web providers retain documents?
Yes. When a web hosting or development relationship ends, files may remain on the former provider's servers, accessible via the original URL. On contract completion, require written confirmation that all business files have been deleted or returned, and check old domain URLs directly.
How often should exposure searches be repeated?
The appropriate frequency depends on how often the organisation publishes documents and how frequently website, cloud or domain changes occur. A quarterly minimum is reasonable for most small businesses. An organisation with active web publishing should check more frequently.
What should a document-publishing policy include?
A publishing policy should define information classification, name who can authorise public publication, require review of metadata and hidden content before upload, specify approved redaction tools, restrict cloud-sharing to named users by default, assign ownership of each system and set a scheduled review process.
Can Microsoft Purview help protect documents?
Microsoft Purview provides data-loss prevention policies, information classification, sensitivity labels and sharing reports for Microsoft 365 environments. These tools can help identify overly permissive sharing in SharePoint and OneDrive and enforce classification-based controls. Verify current capabilities in the Microsoft Purview documentation.
What is data-loss prevention?
Data-loss prevention (DLP) describes policies and tools designed to prevent sensitive information from being shared or published in ways that exceed the organisation's defined controls. DLP tools can detect and block transmission of personal data, financial information or credentials in documents and emails.
Can IT Club help review an exposed document?
Yes. Ask the IT Club Advisor about public website files, cloud-sharing links, document permissions, search indexing, redaction or responding to an accidental data exposure.
Administrator Technical Note
This note covers the technical mechanisms relevant to website administrators, IT managers and data-protection officers assessing and remediating document exposure.
Google crawling and indexing
Googlebot and other Google crawlers discover URLs by following links, reading sitemaps and from previously crawled URLs. When a crawler requests a URL, the server response determines whether the content is indexed. A 200 (OK) response on a publicly accessible file allows indexing. A 401 (Unauthorised) or 403 (Forbidden) response signals that authentication is required and prevents indexing. A 404 (Not Found) or 410 (Gone) response confirms the resource is unavailable and prompts deindexing on the next crawl.
site: and filetype: operator limitations
The site: operator returns a subset of what Google has indexed for a domain — not a complete inventory. Google omits duplicate results, low-quality pages and pages filtered for various reasons. filetype: searches return results matching the specified extension but may miss files served with non-standard content types, files behind query strings or files the crawler has not yet processed. Neither operator can substitute for a direct server-side file inventory.
Google Search Console
For verified domain owners, Google Search Console provides more authoritative indexing information than public searches. The URL Inspection tool shows whether a specific URL is indexed, when it was last crawled and any detected issues. The Page Indexing report lists indexed pages and pages excluded from the index with reasons. The Coverage report identifies pages with errors, warnings and excluded states. These tools require domain verification via a DNS record, HTML meta tag or Google Analytics/Tag Manager.
Removals and outdated-content tools
The Removals tool in Search Console allows verified owners to temporarily suppress a URL from Google Search. The suppression currently lasts approximately six months. After this period, the URL may reappear if still accessible. For permanent removal, the source must return a 404 or 410 status or a valid noindex header, after which a recrawl confirms the removal. The Outdated Content Removal request (available at google.com/webmasters/tools/removals) allows third parties to request removal of cached content that no longer matches the live page, without requiring Search Console ownership.
noindex and X-Robots-Tag
The noindex directive can be applied via the HTML meta robots tag for web pages, or via the X-Robots-Tag HTTP response header for non-HTML files such as PDFs. For PDFs, the X-Robots-Tag header must be set by the server configuration, not within the PDF itself. Google processes noindex instructions after crawling the content. Until the next crawl, an indexed file remains in results. noindex does not restrict direct access — a person who knows the URL can still retrieve the file.
robots.txt
The robots.txt file instructs compliant crawlers not to request specific paths. A path blocked by robots.txt may still be indexed if Google discovers the URL from another source — the address appears in results without content, as a 'known URL' entry. robots.txt does not authenticate or restrict access and should not be treated as a security measure.
Authentication and HTTP status codes
The only reliable technical method for preventing unauthorised access is requiring authentication before serving the file. This returns a 401 or 403 HTTP status to any unauthenticated request, preventing indexing and direct access simultaneously. Options include web-server basic authentication, application-layer authentication, VPN access and Azure AD (Entra ID) or Okta-based conditional access policies for cloud-hosted content.
Sitemaps and public media folders
An XML sitemap submitted to Search Console may list document URLs that were included unintentionally. Audit sitemap content to ensure it does not include non-public files. Many content-management systems create a publicly browsable /wp-content/uploads/ or similar media folder. If directory listing is enabled on the web server, Google can crawl and index every file in that folder. Directory listing should be disabled in web server configuration (e.g. Options -Indexes in Apache, autoindex off in Nginx).
PDF and Office document indexing
Google currently extracts text from publicly accessible PDFs, including body text, form fields and some metadata. For Office documents, Google can process the content of DOCX, XLSX and PPTX files including cell values, text frames and, in some cases, document properties. Not every element is guaranteed to appear in search snippets, but the content enters the search index and can be retrieved via exact-phrase searches. Verify the current extent of extraction in Google's documentation.
Cloud-sharing permissions
SharePoint Online supports several sharing levels: People with existing access, Specific people, People in your organisation and Anyone. The Anyone level produces an anonymous sharing link requiring no authentication and potentially indexable if discovered. The SharePoint admin centre (admin.microsoft.com/SharePoint) and Microsoft Purview compliance portal provide tenant-wide sharing reports. Google Workspace Drive admin reports show external sharing and file counts with public access. Review these reports regularly and configure tenant-level policies to restrict anonymous sharing where not required.
Metadata, redaction and document inspection
Microsoft Office applications provide an Inspect Document (or Check for Issues) function that identifies and optionally removes hidden data, document properties, tracked changes, comments, hidden worksheets and embedded objects. For PDFs, Adobe Acrobat Pro and specialist redaction tools such as Workshare Redact, Litera Transact and PDF-XChange provide certified redaction that removes and overwrites underlying content rather than covering it visually. After redaction, verify the output by attempting to select text in the redacted regions and checking document properties.
Data-loss prevention
Microsoft Purview DLP policies can detect and block sharing of documents containing specified sensitive information types (NI numbers, account numbers, named PII) before they are sent externally or published via SharePoint. Google Workspace DLP (within the Admin console under Rules) provides similar capability for Drive and Gmail. These tools supplement process controls but require configuration and testing to function effectively.
ICO breach assessment
Under UK GDPR Article 33, a personal-data breach likely to result in a risk to individuals' rights and freedoms must be reported to the ICO within 72 hours of the organisation becoming aware. The ICO's breach self-assessment tool at ico.org.uk helps organisations evaluate whether a breach is reportable. Key assessment factors include the type and volume of personal data involved, the period of exposure, whether access logs confirm actual access, the sensitivity of the data and the potential harm to individuals. Where in doubt, legal or data-protection advice should be obtained.
Historic domains and web archives
Even after a domain is taken offline, content may be preserved in web archives such as archive.org (the Wayback Machine) and other public caches. Archive.org allows content owners to submit opt-out requests for their domains. Other search engines (Bing, DuckDuckGo, Yandex) maintain their own indexes and caches independently of Google. A comprehensive review must consider all of these, not only Google Search results.
Operational Heartbeat
Document exposure needs an Operational Heartbeat: domains, public files, cloud-sharing links, permissions, metadata, personal data and unresolved findings should be reviewed rather than assumed to remain secure.
Document exposure changes as staff upload new files, websites are redesigned, old domains remain active, suppliers change, public sharing links are created, permissions drift, documents are duplicated, search engines recrawl content, test systems go live, employees leave, publishing responsibilities change, new personal data appears in documents and redaction processes fail.
A recurring document exposure review should check:
- Current domains — are they all inventoried and reviewed?
- Historic domains — are they still active, and if so, is access controlled?
- Subdomains — are all subdomains inventoried and purposeful?
- Indexed file types — what does a site: filetype: search currently reveal?
- Public document libraries — are published files still appropriate?
- Cloud-sharing links — are any set to 'anyone with the link' or public?
- Website upload folders — is directory listing disabled?
- File owners — does every published file have a named responsible person?
- Metadata — are files being sanitised before publication?
- Redaction — are redacted copies genuinely redacted?
- Personal data — have new documents containing personal data been published?
- Credentials — do any published documents contain passwords, tokens or keys?
- Duplicate copies — are the same files accessible via multiple URLs?
- Removal requests — are outstanding suppressions still appropriate?
- Incident records — have findings from the last review been resolved?
- Supplier access — do former providers still host files?
- Unresolved findings — are actions from previous reviews complete?
- Corrective actions — have process controls been updated?
- Next review date — is it scheduled?
Set a review date. Document exposure is not a one-time check. Files accumulate, permissions change, systems are added and removed and staff turnover means that knowledge of what was published and when is frequently incomplete.
Plain-English Takeaway
Google can index publicly accessible PDFs, Word documents, spreadsheets, presentations and other files. Businesses can use authorised site and file-type searches to identify documents associated with their own domains, but search results are incomplete and removal from Google does not secure the original file. Sensitive content should be deleted, corrected or placed behind proper authentication, with any personal-data exposure assessed through the organisation's incident-response process.
Need the practical steps?
A short, instruction-led version of this topic is available in the Knowledge Centre.
View the Knowledge Centre GuideRelated Articles
Is Your Business Ready for the Vulnerability Patch Wave?
AI-assisted security tools can search large codebases and identify possible software vulnerabilities much faster than traditional manual research alone. This may create a Vulnerability Patch Wave — a sustained increase in security advisories, emergency fixes and updates that organisations must assess, test and deploy faster than before. This article explains what the wave is, why discovery is accelerating, why fixing remains slower and what businesses should do.
Read articleIs Temporary Chat Safe for Sensitive Business Discussions?
ChatGPT Temporary Chat does not appear in normal chat history, does not use or create saved memories and is not used to improve OpenAI's models. However, OpenAI may retain a copy for up to 30 days for safety purposes, Custom Instructions may still apply and the information is still transmitted to and processed by an external service. This article explains what Temporary Chat actually protects and why confidential business information still requires care.
Read articleCan Someone Pretend to Email Your Customers?
A customer receives an email that looks exactly like it came from your business. It asks them to pay an invoice, click a link or reset a password. It did not come from you. This is email spoofing — and DMARC is one of the best tools available to stop it.
Read article