Home Backend Development PHP Tutorial PHPAnalysis Chinese word segmentation practical tutorial_PHP tutorial

PHPAnalysis Chinese word segmentation practical tutorial_PHP tutorial

Jul 13, 2016 am 10:28 AM

PHPAnalysis is a widely used Chinese word segmentation class. It uses reverse matching mode word segmentation, so it is compatible with a wider range of encodings. Its variables and common functions are now explained in detail as follows:

1. More important member variables
$resultType = 1 generated word segmentation result data type (1 is all, 2 is dictionary vocabulary and a single Chinese, Japanese, Korean, simplified and traditional character and English, 3 is dictionary vocabulary and English)
This variable is generally set using the SetResultType($rstype) method.
$notSplitLen = 5 Split the shortest sentence length
$toLower = false Convert all English words to lowercase
$differMax = false Use the maximum split mode to disambiguate bigram words
$unitWord = true Try to merge words (that is, new word recognition)
$differFreq = false Use popular word priority mode for disambiguation
2. List of main member functions
1. public function __construct($source_charset='utf- 8', $target_charset='utf-8', $load_all=true, $source='')
Function description: Constructor
Parameter list: (www.jbxue.com)
$source_charset source String encoding
$target_charset Directory string encoding
$load_all Whether to load the dictionary completely (this parameter has been invalidated)
$source source string
If the input and output are both utf-8, it is actually OK There is no need to use any parameters for initialization, but set the text to be operated through the SetSource method
2. public function SetSource( $source, $source_charset='utf-8', $target_charset='utf-8' )
Function description: Set source string
Parameter list:
$source source string
$source_charset source string encoding
$target_charset directory string encoding
Return value: bool
3 , public function StartAnalysis($optimize=true)
Function description: Start performing word segmentation operation
Parameter list:
$optimize Whether to try to optimize the results after word segmentation
Return value: void
A basic Word segmentation process:
////////////////////////////////////////
$pa = new PhpAnalysis();
$pa->SetSource('String that needs to be segmented');
//Set the segmentation attribute
$pa->resultType = 2;
$pa ->differMax = true;
$pa->StartAnalysis();
//Get the results you want
$pa->GetFinallyIndex();
///// /////////////////////////////////////
4. public function SetResultType( $rstype )
Function description: Setting the type of the returned result
is actually an operation on the member variable $resultType
The value of parameter $rstype is:
1 is all, 2 is dictionary vocabulary and a single Chinese, Japanese, Korean, simplified and traditional character and English, 3 Return value for dictionary words and English
: void
5. public function GetFinallyKeywords( $num = 10 )
Function description: Get the number of specified entries with the highest frequency (usually used to extract document keywords)
Parameter list:
$num = 10 Return number of entries
Return value: Keyword list separated by ","
6. public function GetFinallyResult($spword=' ')
Function description: Get the final word segmentation result
Parameter list:
$spword separator between entries
Return value: string
7. public function GetSimpleResult()
Function description: get Rough segmentation result
Return value: array

(Script Academy www.jbxue.com)
8. public function GetSimpleResultAll()
Function description: Get the rough segmentation result containing attribute information
Attributes (1 Chinese words and sentences, 2 ANSI vocabulary (including Full-width), 3 ANSI punctuation marks (including full-width), 4 numbers (including full-width), 5 Chinese punctuation or unrecognizable characters)
Return value: array
9. public function GetFinallyIndex()
Function description: Get hash index array
Return value: array('word'=>count,...) Sort by frequency of occurrence
10. public function MakeDict($source_file, $target_file='')
function Description: Compile the text file dictionary into a dictionary
Parameter list:
$source_file Source text file
$target_file Target file (if not specified, it is the current dictionary)
Return value: void
11. public function ExportDict($targetfile)
Function description: Export all entries of the current dictionary as text files
Parameter list:
$targetfile target file
Return value: void

www.bkjia.comtruehttp: //www.bkjia.com/PHPjc/812980.htmlTechArticlePHPAnalysis is a widely used Chinese word segmentation class. It uses reverse matching mode word segmentation, so it is compatible with a wider range of encodings and is now A detailed explanation of its variables and commonly used functions is as follows: 1. The more important ones...
Statement of this Website
The content of this article is voluntarily contributed by netizens, and the copyright belongs to the original author. This site does not assume corresponding legal responsibility. If you find any content suspected of plagiarism or infringement, please contact admin@php.cn

Hot AI Tools

Undresser.AI Undress

Undresser.AI Undress

AI-powered app for creating realistic nude photos

AI Clothes Remover

AI Clothes Remover

Online AI tool for removing clothes from photos.

Undress AI Tool

Undress AI Tool

Undress images for free

Clothoff.io

Clothoff.io

AI clothes remover

Video Face Swap

Video Face Swap

Swap faces in any video effortlessly with our completely free AI face swap tool!

Hot Tools

Notepad++7.3.1

Notepad++7.3.1

Easy-to-use and free code editor

SublimeText3 Chinese version

SublimeText3 Chinese version

Chinese version, very easy to use

Zend Studio 13.0.1

Zend Studio 13.0.1

Powerful PHP integrated development environment

Dreamweaver CS6

Dreamweaver CS6

Visual web development tools

SublimeText3 Mac version

SublimeText3 Mac version

God-level code editing software (SublimeText3)

Hot Topics

Java Tutorial
1657
14
PHP Tutorial
1257
29
C# Tutorial
1231
24
How does session hijacking work and how can you mitigate it in PHP? How does session hijacking work and how can you mitigate it in PHP? Apr 06, 2025 am 12:02 AM

Session hijacking can be achieved through the following steps: 1. Obtain the session ID, 2. Use the session ID, 3. Keep the session active. The methods to prevent session hijacking in PHP include: 1. Use the session_regenerate_id() function to regenerate the session ID, 2. Store session data through the database, 3. Ensure that all session data is transmitted through HTTPS.

Explain different error types in PHP (Notice, Warning, Fatal Error, Parse Error). Explain different error types in PHP (Notice, Warning, Fatal Error, Parse Error). Apr 08, 2025 am 12:03 AM

There are four main error types in PHP: 1.Notice: the slightest, will not interrupt the program, such as accessing undefined variables; 2. Warning: serious than Notice, will not terminate the program, such as containing no files; 3. FatalError: the most serious, will terminate the program, such as calling no function; 4. ParseError: syntax error, will prevent the program from being executed, such as forgetting to add the end tag.

PHP and Python: Comparing Two Popular Programming Languages PHP and Python: Comparing Two Popular Programming Languages Apr 14, 2025 am 12:13 AM

PHP and Python each have their own advantages, and choose according to project requirements. 1.PHP is suitable for web development, especially for rapid development and maintenance of websites. 2. Python is suitable for data science, machine learning and artificial intelligence, with concise syntax and suitable for beginners.

What are HTTP request methods (GET, POST, PUT, DELETE, etc.) and when should each be used? What are HTTP request methods (GET, POST, PUT, DELETE, etc.) and when should each be used? Apr 09, 2025 am 12:09 AM

HTTP request methods include GET, POST, PUT and DELETE, which are used to obtain, submit, update and delete resources respectively. 1. The GET method is used to obtain resources and is suitable for read operations. 2. The POST method is used to submit data and is often used to create new resources. 3. The PUT method is used to update resources and is suitable for complete updates. 4. The DELETE method is used to delete resources and is suitable for deletion operations.

Explain secure password hashing in PHP (e.g., password_hash, password_verify). Why not use MD5 or SHA1? Explain secure password hashing in PHP (e.g., password_hash, password_verify). Why not use MD5 or SHA1? Apr 17, 2025 am 12:06 AM

In PHP, password_hash and password_verify functions should be used to implement secure password hashing, and MD5 or SHA1 should not be used. 1) password_hash generates a hash containing salt values ​​to enhance security. 2) Password_verify verify password and ensure security by comparing hash values. 3) MD5 and SHA1 are vulnerable and lack salt values, and are not suitable for modern password security.

PHP: A Key Language for Web Development PHP: A Key Language for Web Development Apr 13, 2025 am 12:08 AM

PHP is a scripting language widely used on the server side, especially suitable for web development. 1.PHP can embed HTML, process HTTP requests and responses, and supports a variety of databases. 2.PHP is used to generate dynamic web content, process form data, access databases, etc., with strong community support and open source resources. 3. PHP is an interpreted language, and the execution process includes lexical analysis, grammatical analysis, compilation and execution. 4.PHP can be combined with MySQL for advanced applications such as user registration systems. 5. When debugging PHP, you can use functions such as error_reporting() and var_dump(). 6. Optimize PHP code to use caching mechanisms, optimize database queries and use built-in functions. 7

PHP in Action: Real-World Examples and Applications PHP in Action: Real-World Examples and Applications Apr 14, 2025 am 12:19 AM

PHP is widely used in e-commerce, content management systems and API development. 1) E-commerce: used for shopping cart function and payment processing. 2) Content management system: used for dynamic content generation and user management. 3) API development: used for RESTful API development and API security. Through performance optimization and best practices, the efficiency and maintainability of PHP applications are improved.

Explain Arrow Functions (short closures) introduced in PHP 7.4. Explain Arrow Functions (short closures) introduced in PHP 7.4. Apr 06, 2025 am 12:01 AM

The arrow function was introduced in PHP7.4 and is a simplified form of short closures. 1) They are defined using the => operator, omitting function and use keywords. 2) The arrow function automatically captures the current scope variable without the use keyword. 3) They are often used in callback functions and short calculations to improve code simplicity and readability.

See all articles