跳到主要内容

实操

How to debug Open Graph tags behind a CDN

When a site sits behind Cloudflare, Vercel or Nginx, three HTML documents exist. Here is how to confirm, layer by layer, what the crawler sees.

更新于 2026年3月17日 · 约 4 分钟

When a site sits behind a CDN, three different HTML documents are in play: the one you wrote, the one the origin serves, and the one the crawler receives. A card that looks right in your editor but wrong in a share is usually a mismatch between two of those. Debugging means comparing them.

The three versions

  1. Origin. What the app returns with no cache in front.
  2. Edge. What the CDN serves — possibly a cached or transformed copy.
  3. Crawler. What the edge returns to facebookexternalhit, Twitterbot, LinkedInBot, Slackbot or Discordbot.

If origin and edge differ, the CDN is caching or rewriting. If edge and crawler differ, a rule is keyed on the user agent.

Step 1 — what the crawler actually gets

Request the live URL as Facebook’s crawler and print only the tags that matter:

curl -sL --compressed \
  -A "facebookexternalhit/1.1 (+http://www.facebook.com/externalhit_uatext.php)" \
  https://example.com/page | grep -iE 'og:|twitter:'

Then run the same request with an ordinary browser user agent:

curl -sL --compressed -A "Mozilla/5.0 (compatible; test)" \
  https://example.com/page | grep -iE 'og:|twitter:'

If the two outputs differ, something is varying on the user agent: a bot rule, a challenge page, or a cache keyed by UA.

Step 2 — check status and headers

A crawler that receives a 403 never sees your tags at all. Confirm it gets a 200 and a sane content type:

curl -sI -A "Twitterbot/1.0" https://example.com/page

Watch for a mitigation header (cf-mitigated), a cache status (x-vercel-cache), or a set-cookie that suggests a challenge. A cache-control: no-store is fine. A bot challenge on the crawler is not.

Step 3 — compare against the origin

Bypass the CDN and hit the origin to see the source of truth:

curl -s -A "facebookexternalhit/1.1" \
  --resolve example.com:443:203.0.113.10 \
  https://example.com/page | grep -iE 'og:'

Replace the IP with your origin. If the origin is correct and the edge is not, the fix is at the edge: purge the cache and check for HTML transforms.

What breaks it, per layer

LayerTypical causeWhere to fix
CloudflareBot Fight Mode or a WAF rule challenging social crawlersAllowlist the crawler agents; verify with the curl commands above
CloudflareCached HTML serving a stale headPurge by URL after every deploy
CloudflareHTML minify or a transformer rewriting the headDisable HTML transforms for the page
VercelISR or a CDN cache holding an old renderRedeploy or revalidate; check x-vercel-cache
Nginxproxy_pass with the wrong HostSend the correct Host header to the upstream
Nginxsub_filter or an optimizer stripping tagsExclude the head from rewriting
AnyA redirect chain in front of the pageResolve to a single hop so the crawler sees the final HTML

Re-scraping per platform

PlatformEntry pointNotes
Facebook / MessengerSharing Debugger → “Scrape Again”Also shows the raw tags it received
LinkedInPost InspectorRe-fetches on demand
X (Twitter)Card ValidatorIntermittent; may require login
SlackNo public debuggerPost in a private channel to test
DiscordNo public debuggerAppend a unique query string to bypass the embed cache

After any tag change, purge the edge cache first, then re-scrape. Re-scraping while the edge still serves the old HTML just caches the old card again.

Common mistakes

Fixing the origin and forgetting the edge. Edit, deploy, then purge — in that order.

Caching HTML edge-side with Vary: User-Agent. Different agents get different cached copies, so a crawler can be served a variant rendered for a browser.

Blocking the crawlers you want. Bot protection is the most common reason a card is fine internally and blank externally. Test with the actual crawler user agent.

Trusting a logged-in browser. If the app personalizes the head for signed-in users, your browser shows tags the crawler never will.

Omitting the user agent in the test. A plain curl uses its own user agent; some CDNs treat it as a bot and serve a different response from both the browser and the social crawler.

Re-scraping before purging. The platform caches whatever the edge served. Purge first or you re-cache the bug.

Where this tool fits

The meta preview debugger fetches a URL once and renders the card for each platform, so you can see in a single pass whether the tags survived the CDN. Use it first, then confirm in each platform’s own debugger and purge the cache.

Frequently asked questions

▸ Why are my Open Graph tags wrong only in production?

Production usually has a CDN in front. The edge may serve a cached copy, transform the HTML, or return a different response to crawler user agents. Compare the origin, the edge and the crawler response to find which layer changes the tags.

▸ How do I see what a social crawler receives?

Fetch the URL with curl and the crawler user agent, for example facebookexternalhit or Twitterbot, then filter the output for og: and twitter: tags. That is the document the platform will parse.

▸ Why does Cloudflare block my share previews?

Bot protection such as Bot Fight Mode or a WAF rule can challenge social crawlers before they read the page. Test with the crawler user agent and allowlist it, or add a rule that skips the challenge for those agents.

▸ Do I need to purge the cache before re-scraping?

Yes. If the platform re-scrapes while the edge still serves the old HTML, it caches the old card again. Purge the edge cache first, then trigger the platform re-scrape.

▸ Why do I get different tags with different user agents?

Either a security rule is keyed on the user agent, or the cache stores a variant per user agent. Compare the responses, then make the head identical for every agent so it does not matter which variant is served.