What is the data cleaning method in Python?
The library needed for data cleaning here is the pandas library. The download method is still running in the terminal: pip install pandas.
First we need to read the data
import pandas as pd data = pd.read_csv(r'E:\PYthon\用户价值分析 RFM模型\data.csv') pd.set_option('display.max_columns', 888) # 大于总列数 pd.set_option('display.width', 1000) print(data.head()) print(data.info())
Line 3 It is to read the data. There is a read function call in the pandas library. The csv format is the fastest to read and write.
Lines 4 and 5 are for displaying all the columns when reading, because if there are many columns, pycharm will hide some of the middle columns, so we add these two lines of code to prevent them from being hidden.
The 6th line displays the table header. We can see what fields there are and the column names.
The 7th line displays the basic information of the table. How much data is in each column and what type of field is it? The data. How much non-empty data is there, so in the first step we can see which basic column has a null value.
Null value processing
After data.info() we can see that most of the data has 541909 rows, so we roughly guess it is Description, The CustomerID column is missing results
# 空值处理 print(data.isnull().sum()) # 空值中和,查看每一列的空值 # 空值删除 data.drop(columns=['Description'], inplace=True) print(data.info()) data.isnull()判断是否为空。data.isnumll().sum()计算空值数量。
Line 5 deletes the null value. Here, delete the null value of the Description column first. Inplace=True means to modify the data. If there is no inplace=True, the data will not be modified. , the print data is still the same as before, or a variable is redefined for assignment.
Since there are relatively few null values in this column, this column of data is not that important to our data analysis, so we choose to delete this entire column.
Our table is used to filter customers, so CustomerID is used as the standard and other columns are forced to be deleted.
# CustomerID有空值 # 删除所有列的空值 data.dropna(inplace=True) # print(data.info()) print(data.isnull().sum()) # 由于CustomerID为必须字段,所以强制删除其他列,以CustomerID为准
Here we first perform type conversion on other fields
Type conversion
# 转换为日期类型 data['InvoiceDate'] = pd.to_datetime(data['InvoiceDate']) # CustomerID 转换为整型 data['CustomerID'] = data['CustomerID'].astype('int') print(data.info())
We have dealt with null values above, and next we deal with abnormal values.
Abnormal value processing
To view the basic data distribution of the table, you can use describe
print(data.describe())
You can see that the minimum value in the data Quantity column is -80995. This column obviously has abnormal values , so this column needs to be filtered for outliers.
Only values greater than 0 are required.
data = data[data['Quantity'] > 0] print(data)
When printed, there are only 397924 lines.
Duplicate value processing
# 查看重复值 print(data[data.duplicated()])
There are 5194 rows of duplicate values. The duplicate values here are completely duplicated, so we can delete them as useless data. .
Delete duplicate values
# 删除重复值 data.drop_duplicates(inplace=True) print(data.info())
Save the original table after deletion, and then check the basic information of the table
It’s still there now There are 392730 pieces of data left. At this step, data cleaning is completed.
The above is the detailed content of What is the data cleaning method in Python?. For more information, please follow other related articles on the PHP Chinese website!

Hot AI Tools

Undresser.AI Undress
AI-powered app for creating realistic nude photos

AI Clothes Remover
Online AI tool for removing clothes from photos.

Undress AI Tool
Undress images for free

Clothoff.io
AI clothes remover

Video Face Swap
Swap faces in any video effortlessly with our completely free AI face swap tool!

Hot Article

Hot Tools

Notepad++7.3.1
Easy-to-use and free code editor

SublimeText3 Chinese version
Chinese version, very easy to use

Zend Studio 13.0.1
Powerful PHP integrated development environment

Dreamweaver CS6
Visual web development tools

SublimeText3 Mac version
God-level code editing software (SublimeText3)

Hot Topics











PHP is mainly procedural programming, but also supports object-oriented programming (OOP); Python supports a variety of paradigms, including OOP, functional and procedural programming. PHP is suitable for web development, and Python is suitable for a variety of applications such as data analysis and machine learning.

PHP is suitable for web development and rapid prototyping, and Python is suitable for data science and machine learning. 1.PHP is used for dynamic web development, with simple syntax and suitable for rapid development. 2. Python has concise syntax, is suitable for multiple fields, and has a strong library ecosystem.

PHP originated in 1994 and was developed by RasmusLerdorf. It was originally used to track website visitors and gradually evolved into a server-side scripting language and was widely used in web development. Python was developed by Guidovan Rossum in the late 1980s and was first released in 1991. It emphasizes code readability and simplicity, and is suitable for scientific computing, data analysis and other fields.

Python is more suitable for beginners, with a smooth learning curve and concise syntax; JavaScript is suitable for front-end development, with a steep learning curve and flexible syntax. 1. Python syntax is intuitive and suitable for data science and back-end development. 2. JavaScript is flexible and widely used in front-end and server-side programming.

To run Python code in Sublime Text, you need to install the Python plug-in first, then create a .py file and write the code, and finally press Ctrl B to run the code, and the output will be displayed in the console.

Golang is better than Python in terms of performance and scalability. 1) Golang's compilation-type characteristics and efficient concurrency model make it perform well in high concurrency scenarios. 2) Python, as an interpreted language, executes slowly, but can optimize performance through tools such as Cython.

Writing code in Visual Studio Code (VSCode) is simple and easy to use. Just install VSCode, create a project, select a language, create a file, write code, save and run it. The advantages of VSCode include cross-platform, free and open source, powerful features, rich extensions, and lightweight and fast.

Running Python code in Notepad requires the Python executable and NppExec plug-in to be installed. After installing Python and adding PATH to it, configure the command "python" and the parameter "{CURRENT_DIRECTORY}{FILE_NAME}" in the NppExec plug-in to run Python code in Notepad through the shortcut key "F6".
