实操
How to debug Open Graph tags behind a CDN
When a site sits behind Cloudflare, Vercel or Nginx, three HTML documents exist. Here is how to confirm, layer by layer, what the crawler sees.
更新于 2026年3月17日 · 约 4 分钟
When a site sits behind a CDN, three different HTML documents are in play: the one you wrote, the one the origin serves, and the one the crawler receives. A card that looks right in your editor but wrong in a share is usually a mismatch between two of those. Debugging means comparing them.
The three versions
- Origin. What the app returns with no cache in front.
- Edge. What the CDN serves — possibly a cached or transformed copy.
- Crawler. What the edge returns to
facebookexternalhit,Twitterbot,LinkedInBot,SlackbotorDiscordbot.
If origin and edge differ, the CDN is caching or rewriting. If edge and crawler differ, a rule is keyed on the user agent.
Step 1 — what the crawler actually gets
Request the live URL as Facebook’s crawler and print only the tags that matter:
curl -sL --compressed \
-A "facebookexternalhit/1.1 (+http://www.facebook.com/externalhit_uatext.php)" \
https://example.com/page | grep -iE 'og:|twitter:'
Then run the same request with an ordinary browser user agent:
curl -sL --compressed -A "Mozilla/5.0 (compatible; test)" \
https://example.com/page | grep -iE 'og:|twitter:'
If the two outputs differ, something is varying on the user agent: a bot rule, a challenge page, or a cache keyed by UA.
Step 2 — check status and headers
A crawler that receives a 403 never sees your tags at all. Confirm it gets a 200 and a sane content type:
curl -sI -A "Twitterbot/1.0" https://example.com/page
Watch for a mitigation header (cf-mitigated), a cache status (x-vercel-cache), or a set-cookie that suggests a challenge. A cache-control: no-store is fine. A bot challenge on the crawler is not.
Step 3 — compare against the origin
Bypass the CDN and hit the origin to see the source of truth:
curl -s -A "facebookexternalhit/1.1" \
--resolve example.com:443:203.0.113.10 \
https://example.com/page | grep -iE 'og:'
Replace the IP with your origin. If the origin is correct and the edge is not, the fix is at the edge: purge the cache and check for HTML transforms.
What breaks it, per layer
| Layer | Typical cause | Where to fix |
|---|---|---|
| Cloudflare | Bot Fight Mode or a WAF rule challenging social crawlers | Allowlist the crawler agents; verify with the curl commands above |
| Cloudflare | Cached HTML serving a stale head | Purge by URL after every deploy |
| Cloudflare | HTML minify or a transformer rewriting the head | Disable HTML transforms for the page |
| Vercel | ISR or a CDN cache holding an old render | Redeploy or revalidate; check x-vercel-cache |
| Nginx | proxy_pass with the wrong Host | Send the correct Host header to the upstream |
| Nginx | sub_filter or an optimizer stripping tags | Exclude the head from rewriting |
| Any | A redirect chain in front of the page | Resolve to a single hop so the crawler sees the final HTML |
Re-scraping per platform
| Platform | Entry point | Notes |
|---|---|---|
| Facebook / Messenger | Sharing Debugger → “Scrape Again” | Also shows the raw tags it received |
| Post Inspector | Re-fetches on demand | |
| X (Twitter) | Card Validator | Intermittent; may require login |
| Slack | No public debugger | Post in a private channel to test |
| Discord | No public debugger | Append a unique query string to bypass the embed cache |
After any tag change, purge the edge cache first, then re-scrape. Re-scraping while the edge still serves the old HTML just caches the old card again.
Common mistakes
Fixing the origin and forgetting the edge. Edit, deploy, then purge — in that order.
Caching HTML edge-side with Vary: User-Agent. Different agents get different cached copies, so a crawler can be served a variant rendered for a browser.
Blocking the crawlers you want. Bot protection is the most common reason a card is fine internally and blank externally. Test with the actual crawler user agent.
Trusting a logged-in browser. If the app personalizes the head for signed-in users, your browser shows tags the crawler never will.
Omitting the user agent in the test. A plain curl uses its own user agent; some CDNs treat it as a bot and serve a different response from both the browser and the social crawler.
Re-scraping before purging. The platform caches whatever the edge served. Purge first or you re-cache the bug.
Where this tool fits
The meta preview debugger fetches a URL once and renders the card for each platform, so you can see in a single pass whether the tags survived the CDN. Use it first, then confirm in each platform’s own debugger and purge the cache.
Frequently asked questions
▸ Why are my Open Graph tags wrong only in production?
Production usually has a CDN in front. The edge may serve a cached copy, transform the HTML, or return a different response to crawler user agents. Compare the origin, the edge and the crawler response to find which layer changes the tags.
▸ How do I see what a social crawler receives?
Fetch the URL with curl and the crawler user agent, for example facebookexternalhit or Twitterbot, then filter the output for og: and twitter: tags. That is the document the platform will parse.
▸ Why does Cloudflare block my share previews?
Bot protection such as Bot Fight Mode or a WAF rule can challenge social crawlers before they read the page. Test with the crawler user agent and allowlist it, or add a rule that skips the challenge for those agents.
▸ Do I need to purge the cache before re-scraping?
Yes. If the platform re-scrapes while the edge still serves the old HTML, it caches the old card again. Purge the edge cache first, then trigger the platform re-scrape.
▸ Why do I get different tags with different user agents?
Either a security rule is keyed on the user agent, or the cache stores a variant per user agent. Compare the responses, then make the head identical for every agent so it does not matter which variant is served.