$accumulator是MongoDB 4.4引入的有状态自定义聚合操作符,需显式定义init/accumulate/merge三部分;而$function是无状态JS函数,不支持跨分片合并。

什么是 $accumulator,它和 $function 有什么区别
$accumulator 是 MongoDB 4.4 引入的聚合阶段操作符,用于在 $group 或 $bucketAuto 中定义有状态的、可序列化的自定义聚合逻辑。它不是 JavaScript 函数调用,而是要求你显式声明初始化(init)、累加(accumulate)和合并(merge)三部分——这点和 $function(MongoDB 4.4+ 的无状态 JS 函数)完全不同。$function 不能维护中间状态,也不支持分片环境下的跨分片合并,而 $accumulator 可以。
常见错误是误以为写个 init + accumulate 就够了,漏掉 merge:一旦集合分片或使用 $bucketAuto,缺少 merge 会导致「accumulator must define a merge expression」错误。
必须写的三部分:init / accumulate / merge 怎么写
每个部分都必须返回一个值(通常是对象),且类型需一致。典型场景是实现带权重的中位数、去重计数、或滑动窗口统计——比如「统计每组中 URL 域名出现次数,但相同域名只计一次/文档」:
{
$accumulator: {
init: function() { return []; },
initArgs: [],
accumulate: function(state, url) {
if (!url || typeof url !== 'string') return state;
const domain = url.replace(/^https?:\/\//, '').split('/')[0];
if (!state.includes(domain)) state.push(domain);
return state;
},
accumulateArgs: ["$url"],
merge: function(state1, state2) {
// 合并两个数组并去重
return [...new Set([...state1, ...state2])];
},
lang: "js"
}
}
-
init必须是函数,不能是字面量(如{ count: 0 }不合法,得写function() { return { count: 0 }; }) -
accumulateArgs是字段路径数组,支持嵌套如["user.profile.email"],不支持表达式(如{"$substr": ["$url", 0, 5]}) -
merge的输入是两个init返回类型的值,不能假设顺序,必须幂等
性能与限制:为什么你的 accumulator 跑得慢或失败
$accumulator 在每个文档上执行 JS,且无法利用索引加速;更关键的是,MongoDB 对 JS 执行有严格沙箱限制:
- 禁止
eval、Function构造器、setTimeout等动态执行或异步 API - 单次
accumulate执行超时默认为 10ms(可通过javascriptTimeout参数调高,但不推荐) - state 大小受限(默认 16MB),若累积大量字符串或对象,容易触发
accumulator state exceeds maximum size - 不支持
db.eval()风格的全局变量,所有状态必须通过state参数传递
实际中,如果需要高频更新或大数据量去重,优先考虑用 $addToSet + $size 替代自定义 accumulator——它底层优化过,更快也更稳。
调试技巧:如何验证 accumulator 行为是否符合预期
直接在 $group 中测试太难定位问题,建议拆解验证:
- 先用
db.collection.aggregate([{$limit: 10}, {$project: {domain: {$replaceAll: {input: "$url", find: "https://", replacement: ""}}}}])检查输入字段是否为空或格式异常 - 把
init/accumulate/merge逻辑复制到 Node.js 里,用真实数据手动跑几轮,确认合并逻辑不丢数据(比如merge([a,b], [b,c])应得[a,b,c]) - 在聚合中临时加
{$addFields: {debug_state: "$myAccumulatorField"}},观察中间 state 结构是否如预期(注意:state 是 BSON 对象,可能被自动转成数组或文档)
最容易被忽略的是分片场景下 merge 的调用时机——它不仅在最终结果合并时触发,也可能在中间 shuffle 阶段多次调用,所以 merge 函数必须能接受任意两个合法 state 的组合,不能依赖“只合并两次”这种假设。

















