
本文介绍如何在 polars 中简洁、高性能地对多个数值列(如 value1/value2/value3)按多个时间窗口(如 "30d"、"90d")和分组键(如 strike/maturity)计算 z-score,避免冗余聚合,充分利用窗口函数与正则列选择。
本文介绍如何在 polars 中简洁、高性能地对多个数值列(如 value1/value2/value3)按多个时间窗口(如 "30d"、"90d")和分组键(如 strike/maturity)计算 z-score,避免冗余聚合,充分利用窗口函数与正则列选择。
在 Polars 中实现带分组的滚动 Z-Score(即 (x - rolling_mean) / rolling_std),关键在于避免显式两阶段聚合(先算均值/标准差再做减法除法),而应直接利用 over() 窗口函数 + 滚动表达式组合,让 Polars 在单次扫描中完成广播计算,既简洁又高效。
以下为推荐的最佳实践方案:
✅ 核心思路:rolling_*().over() + 正则列匹配 + 函数化封装
Polars 的 .over(["strike", "maturity"]) 等价于 pandas 的 groupby(...).transform(...),它确保滚动统计(如 rolling_mean)在每个分组内独立计算,并自动广播对齐到原始行长度,无需 first() 或 agg 后 join。
同时,使用正则列选择 pl.col("^value.*$") 可一次性对所有目标列(value1, value2, value3 等)应用相同逻辑,大幅简化代码。
✅ 基础单窗口示例(30天)
import polars as pl
result = (
df.lazy()
.sort("date", "strike", "maturity")
.set_sorted("date") # 显式声明 date 已排序,加速 rolling 操作
.with_columns(
(
(pl.col("^value.*$")
- pl.col("^value.*$").rolling_mean(window_size="30d", by="date", min_periods=1))
/ pl.col("^value.*$").rolling_std(window_size="30d", by="date", min_periods=1)
)
.over(["strike", "maturity"])
.name.suffix("_zscore_30d")
)
.select("date", "maturity", "strike", pl.col("^.*zscore_30d$"))
.collect()
)⚠️ 注意事项:
- min_periods=1 避免起始窗口因数据不足返回 null;
- 必须调用 .set_sorted("date")(在 .sort() 后),否则 rolling_* 在 by="date" 下性能显著下降;
- name.suffix() 自动为每列添加后缀(如 value1_zscore_30d),无需手动拼接字符串。
✅ 扩展至多窗口:函数化 + 循环生成
为支持 "30d", "90d", "180d" 等多个窗口,建议将逻辑封装为可复用函数:
def rolling_zscore(
expr: pl.Expr,
window_size: str,
by: str = "date",
min_periods: int = 1,
**kwargs
) -> pl.Expr:
"""对表达式计算指定窗口的滚动 Z-Score,支持多列向量化"""
mean = expr.rolling_mean(window_size=window_size, by=by, min_periods=min_periods, **kwargs)
std = expr.rolling_std(window_size=window_size, by=by, min_periods=min_periods, **kwargs)
return ((expr - mean) / std).over(["strike", "maturity"]).name.suffix(f"_zscore_{window_size}")
# 应用多个窗口(并行执行!)
windows = ["30d", "90d", "180d"]
result = (
df.lazy()
.sort("date", "strike", "maturity")
.set_sorted("date")
.with_columns(
rolling_zscore(pl.col("^value.*$"), window_size=w)
for w in windows
)
.select(
"date", "maturity", "strike",
pl.col("^.*zscore_.*$")
)
.collect()
)✅ 该写法优势明显:
- 零冗余计算:每个 rolling_mean/rolling_std 仅执行一次,而非重复聚合;
- 完全并行:.with_columns([...]) 中所有表达式由 Polars 自动并行调度;
- 强可维护性:新增窗口只需追加到 windows 列表;
- 类型安全 & 提前报错:编译期检查列名与表达式合法性。
? 最终输出结构说明
假设原始 df 含列 ["date", "maturity", "strike", "value1", "value2"],执行上述多窗口后,结果将包含:
date | maturity | strike | value1_zscore_30d | value2_zscore_30d | value1_zscore_90d | value2_zscore_90d | ...
? 进阶提示:若需动态指定分组列(如有时按 ["strike"],有时按 ["strike","matu"]),可将 over(...) 参数也作为函数参数传入,进一步提升复用性。
综上,告别冗长的 agg + join 模式,拥抱 rolling_*.over().name.suffix() 组合——这是 Polars 原生、高效、地道的分组滚动标准化实现方式。

















