Three weeks after launch I searched site:patet.xyz on Google.
One result. The homepage.
There were twenty-odd pages on the site by then, all loading fine, all interlinked.
First, when slow is normal
New sites take a while to get indexed. Days to weeks between submitting and showing up is unremarkable, so for the first two weeks I did nothing and waited.
Still one page at week three isn't slow. It's blocked.
Here's the order I checked things in — most obvious to least. It's also genuinely the order I went in.
1. Is robots.txt blocking
The first thing to look at. Mine came from a starter template and I had never read the second line:
User-agent: *
Disallow: /The whole site, disallowed. Might as well not have written it — search engines politely stayed out.
Search Console distinguishes this cleanly: "Blocked by robots.txt" and "Discovered – currently not indexed" are completely different states. The first means you're refusing. The second means it's queued.
I touched on this in section 8 of the previous post, so I'll keep it short: fix it, resubmit, wait.
A week later, still one page.
2. Is there a noindex on the page
I read the source. The <head> was clean.
So look elsewhere. There are three places noindex hides, and only one of them is a <meta> tag:
- The
X-Robots-Tagresponse header. Easiest to miss, because it isn't in the HTML at all. - Injected at build time. Some templates ship a production check that flips it on, or a plugin adds it to everything but the homepage.
Checking from the command line beats reading source in a browser:
curl -sI https://your-domain/some-post | grep -i x-robots-tag
curl -s https://your-domain/some-post | grep -i 'name="robots"'Both empty. Ruled out.
3. Where does canonical point
At this point I was out of obvious ideas, so I dropped an inner page into Search Console's URL Inspection.
The verdict: "Alternate page with proper canonical tag." It believed the real version of that article lived somewhere else.
That somewhere was http://localhost:3000.
I have a siteUrl constant in the code that builds canonical URLs, Open Graph tags, and the absolute URLs in the sitemap. On deploy day I changed the DNS records, changed the environment variables, and forgot to change that constant. So every page's <head> was saying the same thing:
<link rel="canonical" href="http://localhost:3000/blog/xxx"/>What a crawler reads is: the canonical version of all these pages is on a host I can't reach. Of course it indexed none of them.
This one hides well because the page itself is perfectly fine. It loads, the links work, the layout is right. You only catch it by viewing source or reading the canonical report in Search Console.
One related trap: if your canonical points at www while you actually serve the bare domain (or the reverse), you get exactly the same outcome. Buying a domain suggests picking one up front and 301-ing the other.
4. Does the sitemap actually contain URLs
Last one, and the dumbest.
I did submit sitemap.xml, and Search Console did say it was read successfully. What I never did was open it and count.
My sitemap was generated at build time, and the code generating it read from a hand-maintained list of routes. I added pages later and forgot to add them to the list. The result was a perfectly well-formed <urlset> with zero <url> elements inside.
Search Console reports "Success." It does not report "found 0 URLs."
The lesson: "read successfully" doesn't mean "has content." Count it after you submit:
curl -s https://your-domain/sitemap.xml | grep -c "<loc>"How it resolved
The day after all four were fixed, results started appearing. Within a week most pages were in the index.
One thing worth stating plainly: being indexed is not the same as ranking.
Everything above is about making a crawler willing to come in and able to understand what it finds. Where you place in results is content quality and links, which is a different post.
Three things I took away
- Indexing problems are almost always you telling search engines not to index you. robots, noindex, canonical — all three are things you wrote. None of it is a search engine being unfair.
site:is a rough indicator, not a metric. The numbers are unreliable; use the coverage report in Search Console when you want the real state.- Verify every step instead of assuming it. All three of my mistakes were variations of "I was sure I'd configured that."
SEO on this site is the positive checklist — what to do. This post is the negative one — what it looks like when you get it wrong. Read together, they just about cover the path to getting indexed.
And one coincidence worth mentioning: the three things I got wrong — robots, sitemap, canonical — are precisely the three items under "SEO & discoverability" in the launch checklist. Walking that list first would have saved me three weeks.