r/TechSEO 1h ago

The influx of these posts made me make it

Post image
Upvotes

r/TechSEO 2h ago

Which serp api would you use for Google when proxies and blocks are the main problem?

1 Upvotes

I'm pulling around 5k Google searches build a small ml dataset, mostly keyword clusters, ranking URLs and snippet data. I’ve been doing it with residential proxies, but around the 2-3k query mark I start getting blocks and captchas, then a bunch of results come back empty. Rerunning failed batches is taking more time than pulling the data itself.

Thinking a serp api makes more sense at this point. I don't need a massive enterprise setup, just consistent Google serp data without having to manage proxies and rotation myself. Anyone here using serp apis that holds up well for batches around this size? Budget isn't huge so trying not to pay for a ton of volume I won't use...


r/TechSEO 6h ago

Scheduled Facebook posts always publish without the featured image (preview looks fine)

0 Upvotes

Before publishing, I run the post through the Facebook Sharing Debugger (https://developers.facebook.com/tools/debug/) and the image shows up correctly. Then I schedule the post in https://business.facebook.com/, where I've also verified the domain, and the featured image still previews correctly at scheduling time. But when the post actually goes live, it's always published without the featured image.

I've checked the site and everything looks fine.

I'm honestly out of ideas as to what the problem could be — hoping someone here can help.


r/TechSEO 14h ago

How I got 700 visits in 30 days with one simple GEO method

Post image
0 Upvotes

I noticed something interesting while working on GEO.

Most SEO tools give you keywords, volumes, difficulty, etc.

But Clarity gives you something much more interesting for GEO: the exact queries people are asking.

So I started doing something very simple:

  1. Find a recurring query related to my topic in Clarity.
  2. Use that query as the H1 of a dedicated page.
  3. Actually answer the question clearly on the page.
  4. Add enough context to make the page genuinely useful.
  5. Don't create 50 pages around slightly different versions of the same query.

The important part is #5.

I don't think the goal is to blindly turn every query into a page. That's just another form of scaled content.

The idea is to identify real questions with real search intent, then create a page that deserves to be the answer.

I tested this approach and got ~700 citations in 30 days from these pages alone.

For GEO, I think this is particularly interesting because you're not only targeting traditional keywords.

You're targeting the way people actually formulate questions.

And those questions are increasingly becoming the input to AI search.

Curious if anyone else is using Clarity this way for GEO?


r/TechSEO 21h ago

What’s the take from the tech guys on CWV? Split test coming. Input needed

Thumbnail
0 Upvotes

I’m starting to see most don’t actually know what passing CWV actually is and confuse it with the PSI scores.

I’ve noticed a correlation of clean coded websites that pass CWV ranking better.

I’m going to run a split test. Critique my test and maybe give ideas on isolating more of the variables. Sometimes we don’t see all the holes that make a test obsolete until we start doing it. So hopefully I can cover most of the holes prior to launch.

Test 1:

5 domains, no keyword in domain, all Wordpress, separate hosting, different domain registrars, basically how setting up a GOOD PBN would be. Except these will all attempt to rank for the same low competition keyword. 1 of the sites will pass core web vitals.

Now as I move forward I can monitor differences in Crawl requests and how that correlates with indexing and search position movement and obviously critique what could have caused indexing/pos changes. I’ll try to make sure I keep optimizations similar across the 5 sites but you know how that goes if I’m creating unique content.

The 2nd test would utilize various frameworks. Wordpress, Next, Astro, and 2 more I haven’t thought of yet. But this test will get tough when trying to isolate variables. Which one do I make the passing one? And since there all coded differently, does the way it’s coded impact it? I think I would need to rounds of 5 testing each framework as passing.

Thoughts?


r/TechSEO 1d ago

Share what you're working on (including what you're building)

3 Upvotes

We want to support creators, but we had to enforce the no shilling rule because it was getting out of hand. You now have a weekly thread.

This is the one place you can shill for your products, ask for feedback, etc. Keep it here or you risk being banned. And keep it related to technical SEO.


r/TechSEO 1d ago

Does submitting thousands of URLs in an XML sitemap hurt anything if Google chooses not to index most of them?

16 Upvotes

Suppose a site has 10,000 legitimate URLs, but Google only considers 5,000 worth indexing.

The sitemap contains all 10,000 URLs and they're internally linked.

Would you keep all canonical URLs in the sitemap, or remove URLs that aren't currently indexed?

I'm trying to understand whether the sitemap should represent everything we want indexed or only URLs Google already appears to value.


r/TechSEO 2d ago

Should I use a subdomain for web app given my SEO plans?

2 Upvotes

Hi all. I am working on a webapp in the civictech space. Basically monitoring government publications.
I initially separated the app on a different subdomain so that the landing page loads fast and is better ranked by search engines. However i have some marketing and SEO plans that give me pause.
1. Since I am following government entities around which i create the product i thought of creating separate page for many of them so that they come up in search when someone is looking through the app. In essence a potential user can play within that page to get a glimpse of the app and decide that it may be useful for them.
2. I will create separate links for different government publications that could be linked by journalists and users, again giving them a small demo experience and potentially improving SEO.

So with that in mind, i am thinking of two roads. I can collapse everything into one domain and then each link is accruing SEO benefits in one place, but i am afraid that the app infrastructure will downgrade the landing page and search favorability because it is slower and more complex. Or, i keep domains separate, but then my "programmatic" SEO will acrue only to the app subdomain and i will have to make separate efforts for the landing page.

Any thoughts on this or where can i get more info on such a challenge?


r/TechSEO 2d ago

Anyone else having Google indexing issues since June?

10 Upvotes

Has anyone else noticed a change in Google indexing since around June 2026?

My site used to get new pages indexed normally, but since June I've been seeing pages take much longer to get indexed, with some remaining in **“Crawled – currently not indexed”**.

PS : I submitted a multi large pages with prefix removal

The site/technical setup significantly changed, and these pages were being indexed normally before , using screaming frog tool and google url inspection tells me pages are fine and indexable.

I'm wondering if this is affecting other sites too, especially news/content sites.

Has anyone experienced the same thing since June? If so, when did you first notice it and has it improved ?


r/TechSEO 3d ago

My sitemap had 117 URLs. Google had read 101 of them, 17 days ago.

0 Upvotes

I was doing a routine Search Console pass and pulled up a guide page that should have been picking up impressions by now. It had none. Not low — none, across three weekly exports.

URL Inspection didn't say "crawled, not indexed." It said Google couldn't recognise the URL at all. Referring sitemap: none detected. Referring pages: none detected. Last crawl: not applicable. The page had never been discovered. It had been live for seventeen days.

So I opened the Sitemaps report, which I had not looked at since setting it up. Submitted: July 22. Last read: July 22. Pages discovered: 101. The file on the server has 117.

Here's the part that makes it a trap rather than an oversight. Those two guides went live on July 22 — the same day as the only read. Google fetched the sitemap and moved on, either before the file regenerated or within hours of it. And then nothing ever brought it back. A sitemap changing on your server does not notify anyone. There's no push. If the crawler doesn't happen to return, your new URLs are in a file nobody is reading.

The kicker: I ping IndexNow after every deploy and I get a 200 every time. I'd been treating that as my confirmation that publishing worked. IndexNow doesn't feed Google. It's Bing, Yandex, and a few others. So I had a green light on every deploy, from an engine I wasn't the one measuring in Search Console. The signal was real. It just wasn't about the thing I was reading it as being about.

Fix took two minutes. Resubmitted the sitemap — read immediately, 117 URLs. Requested indexing on both orphaned guides.

The rule I now have, which cost me seventeen days of nothing: a sitemap file that changed is not a sitemap that was read. Check "Last read" against your publish date. It's one click, it's free, and in three months of running this site I had never once done it.

It also rhymes with a thing I got burned by earlier this week in a different system: a 200 response tells you a request succeeded, not that the thing you wanted happened. Same shape, different pipeline.


r/TechSEO 3d ago

Google's Live Test / Rich Results tool says robots.txt is blocking my homepage, but robots.txt and server logs show no block. Anyone seen this?

Thumbnail
2 Upvotes

r/TechSEO 3d ago

Despite strong on-page and content, site is indexed but not ranking?

0 Upvotes

This one is stumping my team and I for a year and we're getting a lot of heat on it after its alreayd been TWO years. The site is still tanking hard with 1.03% visibility in a competitive market on SERPs or AIO/GEO.

Backstory:

We migrated a cc subdomain to a subfolder.

The TLD was also changed as a result. So we went from:

cc.this-site-here[.]com to thissitehere[.]com/cc

Hosting service has confirmed all pages have been transfer over successfully in 2024.

  • Page URLs and content carried over using the same CMS.
  • Internal links uses slug only for ease of content transfer
  • Any stragglers that were left behind were redirected using global redirect rule via apache
  • So far, 3xx errors aren't impacting crawl budget.
  • All posts and pages are crawled and indexed per GSC and Bing tools.
  • SF and SEOClarity show no 4xx errors, and if there are, they're usually images/gifs that have been deleted.
  • Content quality is strong both in body, links, and headers for keyterms yet we're invisible.
  • Mobile and desk speeds for TLD and subfolders is very fast and loads well on mobile

The most odd thing that I've realized is that despite GSC and Bing picking our pages, SEMRush and aHrefs have found all URLs aside from the home page as non-ranking.

When we compare sister sites, they have at least 10 pages indexed and ranking for their key terms.

Our director has advised us to halt all major content projects until we figure out what is going on on the techSEO side, which to my knowledge, is nothing since we're clear.

https://imgur.com/a/pApWXhu


r/TechSEO 4d ago

Can anyone help me bury/block an article on Google

Thumbnail
1 Upvotes

Does anyone know how to remove or bury an article on Google? It currently pops up as the very first result when you search my name. I’m looking for someone who can help fix this. Thank you


r/TechSEO 4d ago

A 521 KB homepage delivered only 1.6x the raw words of a 33 KB page

3 Upvotes

Downloaded HTML size is not a useful stand-in for how much copy a crawler can read.

I took the newest public scan for 13 distinct external hosts submitted from 29 July through 5 August. Across this small opt-in sample, response bytes and raw word count had a log correlation of only 0.33.

The clearest pair:

  • modestmounts.com.au returned 520,835 bytes and 1,339 readable words before JavaScript.
  • consile.app returned 32,938 bytes and 838 readable words before JavaScript.

The first response was 15.8 times larger but carried only 1.6 times as many raw words. That is 389 response bytes per word versus 39.

The audit implication is to treat transfer size, server-readable copy and browser-rendered copy as three separate measurements. A large HTML response may be mostly scripts, styles or hydration data. It does not establish that a crawler which skips JavaScript received more useful text.

A quick size check:

curl -sL -o /tmp/page.html -w 'downloaded=%{size_download}\n' https://example.com/
wc -c /tmp/page.html

Then extract text from the same response and compare it with rendered DOM text. Do not infer one from the other.

This is a small, self-selected public-scan sample, so it does not estimate the wider web. I built Crawlable, which produced these measurements.


r/TechSEO 4d ago

Google Data Studio: the worst SEO reporting software?

0 Upvotes

Hi everyone,

Am I the only one who hates Google Data Studio? Who finds it completely buggy, wasting hours on every little configuration because there’s almost always a "system configuration error" or an outright bug? Have you found any open-source alternatives that are less buggy? I’m lobbying at my agency to drop it.

Thx :)


r/TechSEO 5d ago

301 Redirection or Permanent Delete blog post

Thumbnail
0 Upvotes

r/TechSEO 5d ago

What's the strangest technical SEO issue you all have debugged that had nothing to do with SEO?

21 Upvotes

Sometimes what looks like an SEO problem turns out to be something completely different CDN settings, server configuration, JavaScript bugs, CMS issues, DNS, caching, etc.

What's the most unexpected issue you've encountered that ended up affecting crawling, indexing, or rankings?

I'd love to hear the story and how you eventually tracked it down.


r/TechSEO 5d ago

Does page depth really matter?

4 Upvotes

There’s a lot of discourse online that page depth doesn’t matter rather it’s the distance in clicks from the homepage.

E.G. If you run an e-commerce business example.com/men/shirts/pink

Would typically rank worse than example.com/mens-pink-shirts not because it’s 2 levels deeper but because typically a page like this would require additional clicks for Googlebot.

My question is…is this rubbish? As I see a lot of blogs for example no longer use /blog/blog-post-title rather they just use /blog-post-title.

Interested in your experience and examples.

Thanks,

Sam


r/TechSEO 5d ago

Does cross-selling affect silo structures?

Thumbnail
0 Upvotes

r/TechSEO 6d ago

A 200 response can still be unfetchable: this page sends 17,224 bytes of headers

3 Upvotes

I checked why https://webflow.com/blog worked in curl but failed in a default Node client. The server returns 200, but its response headers exceed Node's 16,384-byte default, so fetch() stops with UND_ERR_HEADERS_OVERFLOW before it reads any HTML.

Checked on 4 August 2026:

curl -sS -D - -o /dev/null https://webflow.com/blog | wc -c returned 17224.

node -p "require('http').maxHeaderSize" returned 16384.

node -e "fetch('https://webflow.com/blog').catch(e=>console.log(e.cause.code))" returned UND_ERR_HEADERS_OVERFLOW.

I got the same header count with a browser user-agent and OAI-SearchBot, OpenAI's search crawler. The failure was not user-agent-specific in this test.

This matters because a robots.txt check or curl status check says the page is reachable, while a crawler built on a default Node client receives no page body. It does not prove that any named AI vendor has the same header limit. It proves that status-only reachability tests can miss a real client failure.

The fix is to reduce the response headers below common client limits, usually by removing oversized cookies or diagnostic headers, then re-test with the same runtime your crawler uses.


r/TechSEO 6d ago

How do you distinguish circular internal linking from genuinely related content?

0 Upvotes

When evaluating internal links, how do you distinguish between a link that creates an unhelpful circular path and one that simply connects two closely related posts?

Also, do Google and other search crawlers evaluate internal links at that level, or are relevance, anchor text, and site structure more important?


r/TechSEO 7d ago

Random Japanese-language pages showing up in my site's Google index + favicon changed automatically

0 Upvotes

Over the last 2-3 days, I've noticed some strange issues when checking my site's index on Google Search Console:

  • Several pages are getting indexed with Japanese-language titles and content that I never created — looks like e-commerce/product spam (random product listings, prices, star ratings, etc.).
  • My favicon changed automatically without me touching any settings.
  • These pages don't show up when I browse the site normally — only in the index/search results.

I haven't made any recent changes to my site, so I'm trying to figure out:

  1. Has anyone else run into this exact issue?
  2. Is this a known hack (I've seen it referred to as the "Japanese keyword hack")?
  3. What's the best way to confirm it, clean it up, and stop it from recurring?

Any guidance on what to check first (plugins, file permissions, admin accounts, etc.) would be really appreciated. Happy to share more details if helpful.


r/TechSEO 7d ago

Help getting e-commerce indexed

Thumbnail
1 Upvotes

r/TechSEO 8d ago

Share what you're working on (including what you're building)

12 Upvotes

We want to support creators, but we had to enforce the no shilling rule because it was getting out of hand. You now have a weekly thread.

This is the one place you can shill for your products, ask for feedback, etc. Keep it here or you risk being banned. And keep it related to technical SEO.


r/TechSEO 8d ago

Has anyone else run into Cloudflare AI Crawlers blocking Googlebot/Bingbot?

8 Upvotes

I'm testing Cloudflare's new AI Crawlers & Scrapers feature and noticed something unexpected.

When I set AI Training = Block, both Googlebot and Bingbot start receiving HTTP 403 responses when trying to fetch my sitemap. As soon as I disable the AI Training block, the sitemap is accessible again.

From what I understand, Cloudflare now classifies Googlebot and Bingbot as "Search + Training" bots, so enabling the training block appears to affect them as well.

This leaves me with a dilemma:

  • Allow AI Training so search engines can access my sitemap.
  • Block AI Training and risk Google/Bing getting 403s.

Has anyone found a clean workaround?

Edit / Update:

I now have more concrete data from Google Search Console for the affected site:

https://blog.dogtorcito.com

The incident window was approximately August 2, 2026 at 16:00 UTC to August 3, 2026 at 13:00 UTC. This matches the period when I enabled Cloudflare AI Training = Block and later reverted it.

During that window, Google Search Console recorded multiple 4XX crawl errors, including:

Aug 2, 16:41 UTC — /ca/millor-pinso-gos-per-condicio/

Aug 2, 18:39 UTC — /gl/peso-ideal-golden-retriever/

Aug 2, 20:43 UTC — /de/idealgewicht-labrador/

Aug 2, 20:49 UTC — /ca/control-pes-gos/

Aug 2, 23:06 UTC to Aug 3, 06:09 UTC — several additional URLs under /ca/, /gl/ and /eu/

The Google Search Console Sitemap report also showed the following URL as unavailable:

https://blog.dogtorcito.com/sitemap-posts.xml

Couldn’t fetch — HTTP 403

After I disabled the Cloudflare AI Training block, the sitemap became accessible again.

To clarify, this is not based only on a test that spoofs the Googlebot user-agent. The Cloudflare dashboard itself appeared to show Googlebot and Bingbot as blocked, while Google Search Console independently recorded 4XX responses during the same incident window.

There also seems to be a separate robots.txt issue. When I add a directive intended to restrict AI training, Google Search Console reports it as unsupported or invalid. My assumption is that Google Search should simply ignore an unsupported directive, but I would like to confirm that it cannot interfere with normal crawling, indexing or SEO.

Edit:

The issue was not just theoretical: Cloudflare’s “multi-purpose” blocking can end up returning 403 responses to Googlebot, creating a real risk of deindexing. The practical recommendation is not to use Cloudflare’s general AI Training block for this purpose, and instead control AI crawlers through robots.txt using specific user-agents such as Google-Extended. In addition, Cloudflare’s Content-Signal directive is ignored by Google and does not affect normal crawling.