
本文介绍一种基于 Jsoup 的健壮 HTML 表格解析方案,解决因 中意外多出 导致表头()与数据行()列数不一致、字段错位的问题,确保 date、Status、FileName、RowsProcessed 等字段精准映射。
本文介绍一种基于 jsoup 的健壮 html 表格解析方案,解决因 `
在使用 Cucumber 进行 Web 自动化验收测试时,常借助 TableParser 解析 HTML 表格数据。但当页面 HTML 存在不规范结构——例如某 <tr> 中因前端逻辑错误多渲染了一个 <code><td>(如状态栏动态插入额外单元格),就会导致表头列数(如 4 列:<code>date / Status / FileName / RowsProcessed)与实际数据行列数(如 5 列)不匹配。此时标准解析器易发生“错位映射”:FileName 被赋值为原 Status 的值,RowsProcessed 取到 FileName 的值,而 Status 反成空字符串——严重破坏断言可靠性。
为此,我们设计了 HtmlTableWithHeader 工具类,核心思想是:严格按表头列定义进行截断式映射,忽略冗余单元格,而非强求行列完全对齐。该类基于 Jsoup 实现,具备高容错性与语义清晰性:
- ✅ 自动识别
<th>(表头)与 <code><td>(数据单元格),统一用 <code>CELL_EVALUATOR提取所有合法单元格; - ✅ 提取表头文本生成字段名列表(
headerLabels); - ✅ 遍历每行数据时,仅取前
min(headerLabels.size(), 实际单元格数)个单元格,逐列绑定至对应字段; - ✅ 返回
List<map string>></map>,天然适配 Cucumber DataTable 断言或自定义校验逻辑。
以下是关键实现代码(已精简注释,突出健壮性设计):
public class HtmlTableWithHeader {
private final Element headerRow;
private final List<Element> bodyRows;
private static final Set<String> CELL_TAGS = Set.of("td", "th");
private static final Evaluator CELL_EVALUATOR = new Evaluator() {
@Override
public boolean matches(Element root, Element element) {
return CELL_TAGS.contains(element.tagName().toLowerCase());
}
};
public static HtmlTableWithHeader parse(Element tableElement) {
if (!"table".equalsIgnoreCase(tableElement.tagName())) {
throw new IllegalArgumentException("Expected <table>, got: " + tableElement.tagName());
}
Elements rows = tableElement.getElementsByTag("tr");
Iterator<Element> it = rows.iterator();
Element header = it.hasNext() ? it.next() : null;
return new HtmlTableWithHeader(header, new ArrayList<>(it));
}
public List<Map<String, String>> getRowsAsListOfMaps() {
if (headerRow == null || bodyRows.isEmpty()) return Collections.emptyList();
// 安全提取表头文本(忽略空/空白)
List<String> headers = Collector.collect(CELL_EVALUATOR, headerRow).stream()
.map(Element::text)
.map(String::trim)
.filter(s -> !s.isEmpty())
.collect(toList());
List<Map<String, String>> result = new ArrayList<>();
for (Element row : bodyRows) {
Elements cells = Collector.collect(CELL_EVALUATOR, row);
Map<String, String> rowMap = new LinkedHashMap<>();
int bound = Math.min(headers.size(), cells.size());
for (int i = 0; i < bound; i++) {
rowMap.put(headers.get(i), cells.get(i).text().trim());
}
result.add(rowMap);
}
return result;
}
}使用示例(Cucumber Step Definition 中):
立即学习“前端免费学习笔记(深入)”;
@Then("the report table contains correct status entries")
public void verifyReportTable(Document doc) {
Element table = doc.getElementById("report-table");
HtmlTableWithHeader parser = HtmlTableWithHeader.parse(table);
List<Map<String, String>> rows = parser.getRowsAsListOfMaps();
// 安全断言:即使某行多一个<td>,Status 字段仍准确取值
assertThat(rows).anySatisfy(row ->
assertThat(row).containsEntry("Status", "SUCCESS")
.containsEntry("FileName", "data_2024.csv")
);
}注意事项:
- 该方案不修复 HTML 源码问题,而是构建面向测试的弹性解析层;生产环境仍应推动前端修复冗余
<td>; <li>若需兼容 <code><colspan></colspan>/<rowspan></rowspan>,需扩展单元格坐标映射逻辑(当前版本暂不支持); - 建议搭配
AssertJ或JUnit 5的assertThat进行结构化断言,避免手动遍历比对; - 在 Cucumber 中,可将
getRowsAsListOfMaps()结果直接转换为DataTable或用于DataTable.diff()对比。
通过此方案,你不再受限于 HTML 的“完美结构”,而是以业务语义(字段名)为中心,实现稳定、可维护、抗干扰的表格验证逻辑。



















