robots meta标签必须写在<head>内且需首次解析即存在,动态注入或置于<body>无效;noindex与nofollow作用独立可组合;robots.txt优先级高于meta robots;X-Robots-Tag更可靠但需服务端支持。

robots meta标签该写在什么位置
必须放在 <head> 区域内,且最好紧贴 <title> 后或 <meta charset> 之后。如果写在 <body> 里,所有主流搜索引擎都会忽略它。
常见错误是把它当成普通 HTML 标签随意放置,或者用 JavaScript 动态插入——后者完全无效,因为爬虫不执行 JS,也看不到运行时生成的 <meta name="robots">。
content 属性值怎么选才有效
最常用组合只有几个,别乱拼:index, follow(默认行为,可省略)、noindex, follow、noindex, nofollow、none(等价于 noindex, nofollow)。
nofollow 单独用没意义——爬虫本来就不会把页面当“链接权重传递源”来处理,除非你明确想阻止它跟踪页面上的外链,才加 nofollow;但更推荐用 rel="nofollow" 控制单个链接。
立即学习“前端免费学习笔记(深入)”;
-
noindex:告诉爬虫别收录当前页,但允许抓取页面内容(比如提取 canonical 或跳转逻辑) -
none:等价于noindex, nofollow,语义清晰,建议优先用 - 避免写
noindex, nofollow, noarchive这类扩展值——noarchive是独立指令,需单独写成<meta name="googlebot" content="noarchive">
robots meta 和 robots.txt 冲突时谁生效
robots.txt 控制的是“能否抓取”,meta name="robots" 控制的是“抓取后是否索引”。两者不是互斥关系,而是先后顺序:先看 robots.txt 是否允许访问,再看页面里的 meta 是否允许索引。
典型陷阱:robots.txt 禁止了某路径(如 Disallow: /private/),那即使页面里写了 index, follow,爬虫根本不会去请求那个页面,也就看不到这个 meta 标签。
反过来,如果 robots.txt 允许抓取,但页面含 <meta name="robots" content="noindex">,爬虫会下载页面,解析后放弃索引——此时仍可能在搜索结果中显示“已删除”或“暂无快照”,属于正常行为。
不同搜索引擎对 robots meta 的支持差异
Google、Bing、Yandex 都支持标准 name="robots",但部分指令有专属写法:
- Google 识别
name="googlebot",可用于覆盖全局规则(例如:全站robots.txt允许,但某页用<meta name="googlebot" content="noindex">单独屏蔽) - Bing 不支持
name="bingbot",只认name="robots",所以别写错名字 - 不要写
<meta name="robots" content="all">—— 这不是标准值,部分旧爬虫可能误判为index, follow,但不可靠,直接写index, follow或留空更稳妥
真正容易被忽略的是:移动端页面若用 rel="canonical" 指向 PC 版,而 PC 版又设了 noindex,那移动端页也会被连带排除——因为 canonical 关系意味着“内容主源已拒绝索引”,这个连锁效应常被漏掉。



















