Yii框架需借助第三方库(如smalot/pdfparser)提取PDF中文文本,推荐使用支持UTF-8和CJK字体的纯PHP解析器,注意PDF需内嵌字体且非扫描件;也可调用系统pdftotext提升准确率。

Yii框架本身不直接提供PDF文本提取功能,它是一个Web应用框架,处理PDF需依赖第三方PHP库(如smalot/pdfparser、setasign/fpdi或spipu/html2pdf等)。要从PDF中准确提取含中文的文本内容,关键在于选用支持UTF-8和CJK字体解析的库,并正确配置编码与字体映射。
推荐方案:使用 smalot/pdfparser(兼容中文)
smalot/pdfparser 是轻量、纯PHP实现的PDF解析器,对中文支持较好(需PDF内嵌字体或使用标准14字体且未加密)。它不依赖扩展(如pdftotext或poppler),适合在共享主机或Docker环境中部署。
- 安装(Composer):
composer require smalot/pdfparser - 确保PDF未加密,且中文字符使用可识别的字体(如 Adobe-GB1、Noto Sans CJK、SimSun 等);若PDF是图片型(扫描件),此方法无效,需先OCR(如用 Tesseract + imagick)
- Yii2中建议封装为一个服务类(如
app/services/PdfTextExtractor.php),避免在控制器中直写解析逻辑
Yii2中完整可用代码示例
以下为可在 Yii2 控制器或服务中直接调用的提取函数(已测试支持简体中文PDF文本):
<?php
use yii\base\Component;
use Smalot\PdfParser\Parser;
<p>class PdfTextExtractor extends Component
{
/**</p><ul><li><p>从PDF文件路径提取纯文本(含中文)</p><div class="aritcle_card flexRow">
<div class="artcardd flexRow">
<a class="aritcle_card_img" href="/xiazai/gongju/2519" title="Yii Framework 2.0.51"><img
src="https://img.php.cn/upload/manual/001/503/042/6a6b03d191dbf935.png" alt="Yii Framework 2.0.51" onerror="this.onerror='';this.src='/static/lhimages/moren/morentu.png'" ></a>
<div class="aritcle_card_info flexColumn">
<a href="/xiazai/gongju/2519" title="Yii Framework 2.0.51">Yii Framework 2.0.51</a>
<p>Yii Framework 2.0.51 官方 Basic 应用模板,适合旧项目兼容、升级验证和开发测试。</p>
</div>
<a href="/xiazai/gongju/2519" title="Yii Framework 2.0.51" class="aritcle_card_btn flexRow flexcenter"><b></b><span>下载</span> </a>
</div>
</div></li><li><p>@param string $filePath 本地PDF绝对路径(如 '@runtime/uploads/sample.pdf')</p></li><li><p>@return string 提取的文本内容,失败返回空字符串
*/
public function extractText($filePath)
{
if (!is_file($filePath) || !is_readable($filePath)) {
return '';
}</p><p>try {
$parser = new Parser();
$pdf = $parser->parseFile($filePath);
$text = $pdf->getText(); // 自动处理多页、字体编码、换行等</p><pre class="brush:php;toolbar:false;"> // 可选:清理多余空白(保留段落换行,压缩连续空格/制表符)
$text = preg_replace('/[ \t]+/', ' ', $text);
$text = preg_replace('/\n\s*\n/', "\n\n", $text);
$text = trim($text);
return $text;} catch (\Exception $e) { \Yii::error('PDF text extraction failed: ' . $e->getMessage(), CLASS); return ''; } } }
在控制器中调用:
use app\services\PdfTextExtractor;
</li></ul><p>public function actionExtract()
{
$filePath = \Yii::getAlias('@app/uploads/demo.pdf');
$extractor = new PdfTextExtractor();
$content = $extractor->extractText($filePath);</p><pre class="brush:php;toolbar:false;">if (empty($content)) {
throw new \yii\web\NotFoundHttpException('无法提取PDF文本,请检查文件路径或格式');
}
return $this->asJson(['text' => $content]);}
常见问题与优化建议
-
乱码或中文缺失:多数因PDF使用自定义字体但未嵌入字形映射。可尝试用
pdfinfo $file查看是否含Language: zh-CN或字体列表;也可改用poppler-utils的pdftotext -enc UTF-8(需服务器支持),再用exec()调用(注意安全过滤) -
性能优化:大PDF(>50页)建议异步处理(如通过 Yii2 的 Queue + Redis),避免超时;可分页提取:
$pdf->getPages()[0]->getText() -
Yii3 用户注意:命名空间改为
Smalot\PdfParser\Parser不变,但需确认 Composer autoload 正确;推荐搭配yiisoft/files处理文件流
替代方案:调用系统 pdftotext(更稳定,需环境支持)
若服务器已安装 poppler-utils(Ubuntu/Debian:`apt install poppler-utils`),可获得更高准确率:
public function extractWithPdftotext($filePath)
{
if (!is_executable('/usr/bin/pdftotext')) {
return '';
}
$escapedPath = escapeshellarg($filePath);
$output = [];
exec("/usr/bin/pdftotext -enc UTF-8 {$escapedPath} - 2>/dev/null", $output, $returnCode);
return $returnCode === 0 ? implode("\n", $output) : '';
}
该方式对含中文字体映射的PDF兼容性更强,且自动处理横向排版、表格结构等,适合生产环境关键业务。


















