
Unwanted Pages Indexed in Google: Find and Fix Them
In September 2026 the person taking over a dental clinic's website found about 60,000 unwanted URLs indexed in Google. Here is how to tell hacked pages from duplicates your own site creates, and the fix Google documents for each.
- 01Pages a hacker added are the harmful kind: Google says injected pages "could harm your site's visitors or your site's performance in search results". Close the hole first, then delete them and return a 404 or 410.
- 02Spam on your own site search pages can be flagged as hacked, so put a noindex on every search results page. Duplicates from filters, tracking links and www or https copies are common and do not break Google's spam policies, and canonical tags or 301 redirects fix them.
- 03robots.txt does not take pages out of Google, and the Removals tool hides URLs for about six months without fixing them, so use it only for urgent cases alongside the permanent fix.
Thousands of unwanted pages in Google are a real problem when a hacker added them or spammers are abusing your site search, and mostly clutter when your own filters, tracking links or duplicate addresses created them. Hacked pages need the security hole closed and a 404 or 410 on every fake address, abused search pages need a noindex, duplicates need a canonical tag or a redirect, and Google's Removals tool only hides any of them for about six months.
In late September 2026, someone who had recently taken over the website of a dental clinic in their group posted on Reddit's r/SEO that Google Search Console showed around 60,000 indexed URLs and roughly 100,000 more as not indexed, for a site with "only a relatively small number of actual pages". A huge number, in the poster's words, were "gambling/casino pages and other random URLs", and they were found when the person taking over the site got access to Search Console. The replies, which we read on 3 October 2026, split between returning a 410 for the spam URLs, redirecting them to an HTML sitemap page, and finding out how the attacker got in first. We have not seen the clinic's site, so we can't say which cause it had. Google's own documentation answers each of those replies, and we quote it below.
Do unwanted pages in Google hurt your site?
That depends on who made them. Pages a hacker added are the harmful kind. Google's spam policies call this page injection, where "hackers are able to add new pages to your site that contain spammy or malicious content", and warn that "these newly-created pages could harm your site's visitors or your site's performance in search results." Once Google detects a hack it appears in the Security issues report, and affected pages "can appear with a warning label in search results".
Pages your own site generated are a smaller problem. Google's canonicalization guide says "Some duplicate content on a site is normal and it's not a violation of Google's spam policies." The costs are practical. Crawling time goes on the copies, and one page spread over many URLs "may make it harder for you to track how your content performs in search results."
One case sits in between. If your site search prints whatever a visitor types, spammers can search your site for casino or pharmacy terms with a phone number attached, link to those results pages and get them indexed under your domain. Google's John Mueller described this on the Search Off the Record podcast on 30 July 2026: "it's not so much that someone has hacked your website to do this, because your website is doing that freely," yet "when we see that happen, we might flag that as hacked." Google may also be slow to spot it. "Maybe it takes a month," Mueller said (transcript), and in that month a customer who looks up your business can see the spam too.
How do you find pages you never made?
Start in Search Console. A site: search on Google only shows a sample, and Google's site: operator page says it "doesn't necessarily return all the URLs that are indexed".
- Compare two numbers. Open the Page indexing report and set the Indexed count beside the number of pages you know your site has. Google's help says "If your site has fewer than 500 pages, you probably don't need to use this report", but this comparison takes a minute and catches exactly this problem. As an illustration, a 40-page clinic site showing 6,000 indexed URLs has unwanted pages.
- Read the addresses. "View data about indexed pages" lists up to 1,000 example URLs. Switch the report's filter from "All known pages" to "Unsubmitted pages only" to see URLs Google knows about that are missing from your sitemaps, then export them.
- See what searchers see. In the Performance report, sort the Pages and Queries tabs by impressions. URLs you don't recognise, and casino, pharmacy or foreign-language searches, are unwanted pages people are already being shown.
- Search for spam words. Google's own example for site owners is
site:example.com viagra casino, which it says "helps with identifying and monitoring spam problems on your site." Try your domain with words like casino, slot, pharmacy and loan, on both the www and non-www versions, because Google says they don't return the same results.
If any of this turns up spam, check three more places before deleting anything. The Security issues report shows whether Google has flagged a hack. Users and permissions should list only people you know, because in Google's guide to one common spam hack "the hacker will typically add themselves as a property owner in Search Console." The Sitemaps report should list only sitemaps you submitted, since "Hackers often modify your sitemap or add new sitemaps to get their URLs indexed more quickly." Keep the spam URLs out of your browser too. Google says "Avoid using a browser to directly view infected pages on your site" and points you to the URL Inspection tool, which shows the page as Google saw it.
Where do unwanted pages come from?
Sort the exported URLs by pattern and match each pattern to one of these groups. The URLs in the first column are illustrations.
| What the URLs look like | Likely cause | Fix |
|---|---|---|
New folders or random file names with casino, pharmacy or foreign-language text, such as example.com/ltjmnjp/341.html | Pages added by a hacker | Close the hole, delete the pages, return 404 or 410 |
Your search results page carrying spam words, such as example.com/?s=casino+bonus+call | Your site search being abused | noindex on every search results page |
One page repeated with ?color=, ?size= or ?sort= | Filters and sorting | Canonical tag to the unfiltered page, and a 404 for filters with no results |
One page repeated with ?utm_source=, ?gclid= or ?sessionid= | Tracking and session parameters | Canonical tag to the clean URL, and internal links to the clean URL |
| http and https, www and non-www, or a staging copy | Hosting and CMS duplicates | 301 redirect to one version, and a password on staging |
Pages a hacker added
Close the way in first, or the pages can come back. Google's list of how spammers get into sites starts with compromised passwords, missed security updates, and insecure themes and plugins, so change every password, update the CMS and remove plugins nobody uses.
Then delete the pages. Google's cleanup guide says "If you just delete the pages and then configure your server to return a 404 status code, the pages will naturally fall out of Google's index with time." That answers the Reddit thread. Google's steps are to fix the vulnerability, delete the hacker's pages and serve a 404, and a 410 does the same job because Google says "All 4xx errors, except 429, are treated the same" (HTTP status codes). If the fake URLs share a folder or a naming pattern, your developer can return a 410 for all of them with one server rule. When the site is clean, request a review in the Security issues report if Google flagged it, which Google says takes "from a few days to a few weeks".
If the site collects personal data, through appointment or enquiry forms for example, find out whether the attacker could reach it before you start deleting files. Google's cleanup guide says to consider "any business, regulatory, or legal responsibilities before you begin cleaning your site or deleting any files." In Singapore, the PDPC's guide on managing and notifying data breaches requires an organisation to assess whether a breach is notifiable within 30 calendar days and, once it is, to notify the Commission no later than three calendar days after that. A breach involving the personal data of 500 or more people is notifiable on scale alone. SingCERT, the Singapore Cyber Emergency Response Team, responds to cybersecurity incidents in Singapore and can be reached at singcert@csa.gov.sg.
Your site search, abused
Mueller's advice was to "take care of that at the source", either by disallowing your search pages in robots.txt or by putting a noindex tag on every search results page. Google's URL structure guide also suggests blocking "URLs that generate search results". If spam search pages are already in Google, add the noindex first and leave them crawlable, because Google has to fetch a page to see the tag: "For the noindex rule to be effective, the page or resource must not be blocked by a robots.txt file" (noindex). Once they have dropped out, a robots.txt rule can stop the crawling too. Don't answer the spam with a server error. Mueller said a 500 makes Google "reduce crawling overall. For the whole website".
Filters, sorting and tracking links
Google's indexing help names this cause directly: "Your site has a large number of duplicate pages, probably because it uses parameters to filter or sort a common collection (for example: type=dress or color=green or sort=price)." Tracking links do the same, and Google's own example of a duplicate URL is a product page reached through an address ending in ?gclid=ABCD.
The fix is a canonical tag on every copy, pointing at the clean page. Most sites we have checked already carry the tag: in our scan of 71 Singapore business websites in July and August 2026, 69 had one in the HTML of the page we checked. What matters on a filter or tracking URL is where that tag points, and Google treats the tag as a preference, saying it "may choose a different page as canonical than you do". URL Inspection shows both the canonical you declared and the one Google picked. Google prefers canonical tags to noindex here, because noindex "will completely block the page from Search," and for filters it adds: "Return an HTTP 404 status code when a filter combination doesn't return results" (faceted navigation).
Google says these copies are labelled duplicate or alternate in the Page indexing report, and that "Having a page marked duplicate or alternate is usually a good thing; it means that we've found the canonical page and indexed it." The number to act on is the indexed one. URLs in the not-indexed pile, like the clinic's 100,000, are already out of search results.
Copies made by your hosting or CMS
Google's canonicalization guide lists "the HTTP and HTTPS versions of a site" and "the demo version of the site is accidentally left accessible to crawlers" among the usual sources of duplicates. Pick one version of your address and send the others to it with permanent redirects, which Google recommends "when you want to get rid of existing duplicate pages", and put a password on any staging or demo copy. Our website migration SEO checklist covers staging protection and redirect maps in more detail. Old addresses of pages you deleted can be left alone, since Google says "404 responses are not necessarily a problem, if the page has been removed without any replacement." A page that moved should 301 to its new address.
Can robots.txt or the Removals tool get rid of them?
Neither removes a page for good. Google says robots.txt "is not a mechanism for keeping a web page out of Google" (robots.txt introduction). A blocked URL can still be indexed through links from elsewhere and turns up in the Page indexing report as "Indexed, though blocked by robots.txt", where Google's advice is to "remove the robots.txt block and use 'noindex'". A block also slows the cleanup, because "Blocked URLs, however, will stay part of your crawl queue much longer" (crawl budget). Keep robots.txt for stopping Google from crawling sections you never want crawled, such as site search and filters, once they are out of the index.
The Removals tool is for emergencies, such as spam showing up when customers search for your business. Google says it takes a page out of results "within a day" (remove a page), each request "lasts only about six months", and the "Remove all URLs with this prefix" option covers a whole folder (Removals tool help). It hides URLs without fixing anything, since it "does not prevent Google from crawling your page, only from showing it in Search results." Keep the tool away from duplicates of your real pages. Google warns it "could remove all versions (http/https and www/non-www) of a URL", which takes your real page out along with the copy.
How long until Google drops them?
It takes weeks to months, depending on the fix. Google removes a 404 or 410 URL when it next crawls it, because "URLs that are already indexed and return a 4xx status code are removed from the index." A noindex only works once Google recrawls the page, and Google says "Depending on the importance of the page on the internet, it may take months for Googlebot to revisit a page." Ask for a recrawl of the URLs that matter most with URL Inspection. After you mark an issue as fixed in the Page indexing report, "Validation typically takes up to about two weeks". Check the Indexed count again a month after the fix.
How do you stop it happening again?
Own your Search Console. The clinic's spam was found when the person taking over the site opened Search Console, so set the property up in your company's name and add whoever runs the site as a user, one of the questions to ask an SEO agency before you sign. Then watch the Indexed count. Google's help says you should see "a gradually increasing count of indexed pages as your site grows", so a sudden jump is your early warning, and the count belongs in what an SEO agency should show you every month. Block your site search from the start with a noindex or a robots.txt rule. Mueller's advice to people who build sites for others was that "just putting that in robots.txt by default makes it so much easier." And keep the CMS, theme and plugins updated, since Google lists missed security updates among the top ways spammers get in.
Frequently asked questions
Do unwanted indexed pages hurt SEO?
Pages a hacker added can. Google says injected pages "could harm your site's visitors or your site's performance in search results", and a site flagged for hacking can carry a warning label in results. Spam on your own site search pages can be flagged as hacked too. Duplicates from filters, tracking links or www and https copies are a smaller problem: Google says some duplicate content is normal and does not break its spam policies.
Should spam URLs return a 404 or a 410?
Either works. Google says all 4xx errors except 429 are treated the same, and it removes indexed URLs that return them when it next crawls them. What matters more is that the pages are gone from the server and the hole that created them is closed, or new ones appear.
Can I use robots.txt to remove pages from Google?
No. Google says robots.txt "is not a mechanism for keeping a web page out of Google." A blocked URL can stay indexed through links from elsewhere, and the block stops Google from seeing a noindex tag. Use a noindex or a 404 to remove pages, and robots.txt only to stop crawling once they are gone.
How long does Google take to drop removed pages?
Google drops a 404 or 410 URL when it next crawls it, and for pages it rarely visits Google says that can take months. The Removals tool hides URLs within about a day, but only for about six months, so pair it with a permanent fix.
Will a sitemap with only my real pages fix it?
No. Google says it "can index any URL that it finds unless you include a noindex directive on the page", whether or not that URL is in your sitemap, and its Sitemaps report help says there is "no guarantee" that a URL in a sitemap will be crawled or indexed. In the clinic's case, Search Console accepted the sitemap but reported 0 discovered pages, which means Google read no page URLs from the file, so check that yours lists your real pages before you resubmit it.
We post findings like this on LinkedIn.