How to solve the problem of count distinct multiple columns in mysql
The reproduced test database is as follows:
CREATE TABLE `test_distinct` ( `id` int(11) NOT NULL AUTO_INCREMENT, `a` varchar(50) CHARACTER SET utf8 DEFAULT NULL, `b` varchar(50) CHARACTER SET utf8 DEFAULT NULL, PRIMARY KEY (`id`) ) ENGINE=InnoDB AUTO_INCREMENT=1 DEFAULT CHARSET=latin1;
The test data in the table is as follows. Now we need to count the number of columns after deduplication of these three columns.
Problem Analysis
My friend gave me four query statements to locate the problem
SELECT COUNT(*) AS cnt FROM test_distinct; SELECT COUNT(DISTINCT id, a, b) as cnt FROM test_distinct; SELECT id, a, b, COUNT(*) AS cnt FROM test_distinct GROUP BY id, a, b HAVING cnt > 1; SELECT l.id AS l_id, l.a AS l_a, l.b AS l_b, r.id AS r_id, r.a AS r_a, r.b AS r_b FROM test_distinct l LEFT JOIN test_distinct r ON l.id = r.id AND l.a = r.a AND l.b = r.b WHERE r.id is NULL or r.id = 'null';
The query results are as follows:
Notice! ! ! From the test data, we can quickly guess where the problem lies, but it turns out that there are more than 30,000 pieces of data in the table, and it is impossible to view the data with the naked eye.
There are two counterintuitive points in the above query results:
The second piece of data is missing after deduplication statistics, but the result of the third piece of data shows There is no identical data.
When using the same table to do a left outer connection, the driving table has data, but the driven table is empty.
Let’s look at the second question first. The official document has the following explanation:
When using the ON clause, the conditions it contains The expression is the same as that used in the WHERE clause. A common situation is to use the ON clause to specify the join conditions of the table, and use the WHERE clause to limit the rows included in the result set.
If there are no matching rows in the right table for the conditions in the ON or USING part of the LEFT JOIN, then the right table uses all columns set to NULL.
You cannot use arithmetic comparison operators (such as =, < or <>) to compare NULL.
SELECT NULL = NULL; SELECT NULL IS NULL;
So the second problem is that the result of NULL=NULL is always False, which results in the two rows originally Equal data results are not equal.
But this does not solve the first problem: why a piece of data disappeared after deduplication. However, we can guess that the missing data is probably related to the NULL value.
We separate the two operations of count and distinct:
SELECT COUNT(*) as cnt FROM (SELECT DISTINCT id, a, b FROM test_distinct) as tmp;
Huh? The result is correct, which means that the query plan generated by count(distinct expr)
may be different from what we imagined. It is not to remove duplicates first and then count. Use explain to analyze the query plan of the two statements. As shown below:
As you can see from the table, the mysql execution engine directly counts count(distinct expr)
As a query, check the official documentation:
Solution
The problem has finally been clarified. There are two ways to solve this problem. The first is to remove duplicates first and then count. The second is to use the IFNULL()
function:
SELECT COUNT(DISTINCT id, a, IFNULL(b, '0')) as cnt FROM test_distinct;
In addition, count( )Use:
SELECT id, a, b, COUNT(*) FROM test_distinct GROUP BY id, a, b; SELECT id, a, b, COUNT(b) FROM test_distinct GROUP BY id, a, b;
Knowledge point
You cannot use arithmetic comparison operators (such as =, ) to compare null values;
count(distinct expr) returns the number of distinct and non-empty rows in the expr column;
COUNT() has two distinct uses: it can be used to count the number of values in a column, or it can be used to count the number of rows. When counting column values, the column value is required to be non-empty (NULL is not counted). When a column or expression is specified in parentheses of the COUNT() function, the function counts the number of results that have a value in the expression. Another function of COUNT() is to count the number of rows in the result set. When MySQL confirms that the expression value within the parentheses cannot be empty, it is actually counting the number of rows. The simplest thing is when we use COUNT(). In this case, the wildcard does not expand to all columns as we guessed. In fact, it will ignore all columns and directly count all rows - "High-Performance MySQL";
In InnoDB, SELECT COUNT(*) and SELECT COUNT(1) are processed in the same way, and there is no performance difference.
The above is the detailed content of How to solve the problem of count distinct multiple columns in mysql. For more information, please follow other related articles on the PHP Chinese website!

Hot AI Tools

Undresser.AI Undress
AI-powered app for creating realistic nude photos

AI Clothes Remover
Online AI tool for removing clothes from photos.

Undress AI Tool
Undress images for free

Clothoff.io
AI clothes remover

Video Face Swap
Swap faces in any video effortlessly with our completely free AI face swap tool!

Hot Article

Hot Tools

Notepad++7.3.1
Easy-to-use and free code editor

SublimeText3 Chinese version
Chinese version, very easy to use

Zend Studio 13.0.1
Powerful PHP integrated development environment

Dreamweaver CS6
Visual web development tools

SublimeText3 Mac version
God-level code editing software (SublimeText3)

Hot Topics

MySQL is an open source relational database management system, mainly used to store and retrieve data quickly and reliably. Its working principle includes client requests, query resolution, execution of queries and return results. Examples of usage include creating tables, inserting and querying data, and advanced features such as JOIN operations. Common errors involve SQL syntax, data types, and permissions, and optimization suggestions include the use of indexes, optimized queries, and partitioning of tables.

MySQL's position in databases and programming is very important. It is an open source relational database management system that is widely used in various application scenarios. 1) MySQL provides efficient data storage, organization and retrieval functions, supporting Web, mobile and enterprise-level systems. 2) It uses a client-server architecture, supports multiple storage engines and index optimization. 3) Basic usages include creating tables and inserting data, and advanced usages involve multi-table JOINs and complex queries. 4) Frequently asked questions such as SQL syntax errors and performance issues can be debugged through the EXPLAIN command and slow query log. 5) Performance optimization methods include rational use of indexes, optimized query and use of caches. Best practices include using transactions and PreparedStatemen

MySQL is chosen for its performance, reliability, ease of use, and community support. 1.MySQL provides efficient data storage and retrieval functions, supporting multiple data types and advanced query operations. 2. Adopt client-server architecture and multiple storage engines to support transaction and query optimization. 3. Easy to use, supports a variety of operating systems and programming languages. 4. Have strong community support and provide rich resources and solutions.

Apache connects to a database requires the following steps: Install the database driver. Configure the web.xml file to create a connection pool. Create a JDBC data source and specify the connection settings. Use the JDBC API to access the database from Java code, including getting connections, creating statements, binding parameters, executing queries or updates, and processing results.

The process of starting MySQL in Docker consists of the following steps: Pull the MySQL image to create and start the container, set the root user password, and map the port verification connection Create the database and the user grants all permissions to the database

The main role of MySQL in web applications is to store and manage data. 1.MySQL efficiently processes user information, product catalogs, transaction records and other data. 2. Through SQL query, developers can extract information from the database to generate dynamic content. 3.MySQL works based on the client-server model to ensure acceptable query speed.

The key to installing MySQL elegantly is to add the official MySQL repository. The specific steps are as follows: Download the MySQL official GPG key to prevent phishing attacks. Add MySQL repository file: rpm -Uvh https://dev.mysql.com/get/mysql80-community-release-el7-3.noarch.rpm Update yum repository cache: yum update installation MySQL: yum install mysql-server startup MySQL service: systemctl start mysqld set up booting

Laravel is a PHP framework for easy building of web applications. It provides a range of powerful features including: Installation: Install the Laravel CLI globally with Composer and create applications in the project directory. Routing: Define the relationship between the URL and the handler in routes/web.php. View: Create a view in resources/views to render the application's interface. Database Integration: Provides out-of-the-box integration with databases such as MySQL and uses migration to create and modify tables. Model and Controller: The model represents the database entity and the controller processes HTTP requests.
