在 Polars 中,可通过 pl.when().then() 结合 pl.int_range() 生成行索引条件,实现对前 n 行的精准赋值(如设为 null 或自定义值),无需转为 pandas 或使用循环,兼具性能与表达力。
在 polars 中,可通过 `pl.when().then()` 结合 `pl.int_range()` 生成行索引条件,实现对前 n 行的精准赋值(如设为 null 或自定义值),无需转为 pandas 或使用循环,兼具性能与表达力。
Polars 作为高性能 DataFrame 库,不支持类似 Pandas 的 .iloc 原地切片赋值(如 df.iloc[0:5] = np.nan),但提供了更函数式、惰性求值的替代方案。核心思路是:基于行号构造布尔掩码,对满足条件的行返回新值,否则保留原值。
以下为完整示例,将前 5 行的 "test" 列替换为 null(即 Polars 中的 None):
import polars as pl
import numpy as np
# 构造初始 DataFrame
df = pl.DataFrame({"test": np.arange(0, 10)})
# 替换前 5 行为 null(等价于 Pandas 中 df.iloc[0:5] = np.nan)
df = df.with_columns(
pl.when(pl.int_range(0, pl.len()) < 5)
.then(None) # 或 pl.lit(None),显式更佳
.otherwise(pl.col("test"))
.alias("test")
)
print(df)输出:
shape: (10, 1) ┌──────┐ │ test │ │ --- │ │ i64 │ ╞══════╡ │ null │ │ null │ │ null │ │ null │ │ null │ │ 5 │ │ 6 │ │ 7 │ │ 8 │ │ 9 │ └──────┘
✅ 关键说明:
- pl.int_range(0, pl.len()) 动态生成从 0 到 n_rows-1 的整数序列(每行对应一个索引);
- pl.when(...).then(...).otherwise(...) 是 Polars 的条件列操作,完全向量化且可优化;
- None 在 Polars 中自动映射为该列对应类型的 null 值(如 i64 列中为 null);
- 若需替换为非 null 值(如 -999),将 then(None) 改为 then(pl.lit(-999)) 即可。
⚠️ 注意事项:
- 不要尝试 df["test"][0:5] = None —— Polars 的列是不可变的,此操作会报错或静默失败;
- pl.len() 返回当前 DataFrame 行数,确保在 with_columns 中动态适配数据长度;
- 如需批量处理多列,可对每列分别应用相同逻辑,或使用 pl.all() + pl.struct 进阶组合(但通常按列显式处理更清晰)。
该方法完全符合 Polars 的声明式编程范式,在保持代码简洁的同时,充分利用其底层 Arrow 引擎的高效执行能力。

















