I have a table with the following columns: <pre class="prettyprint"><code>URL_ID URL_ADDR URL_Time </code></pre> I want to remove duplicates on the <code>URL_ADDR</code> column using a MySQL query. Is it possible to do such a thing without using any programming?

Consider the following test case: <pre class="prettyprint"><code>CREATE TABLE mytb (url_id int, url_addr varchar(100)); INSERT INTO mytb VALUES (1, 'www.google.com'); INSERT INTO mytb VALUES (2, 'www.microsoft.com'); INSERT INTO mytb VALUES (3, 'www.apple.com'); INSERT INTO mytb VALUES (4, 'www.google.com'); INSERT INTO mytb VALUES (5, 'www.cnn.com'); INSERT INTO mytb VALUES (6, 'www.apple.com'); </code></pre> Where our test table now contains: <pre class="prettyprint"><code>SELECT * FROM mytb; +--------+-------------------+ | url_id | url_addr | +--------+-------------------+ | 1 | www.google.com | | 2 | www.microsoft.com | | 3 | www.apple.com | | 4 | www.google.com | | 5 | www.cnn.com | | 6 | www.apple.com | +--------+-------------------+ 5 rows in set (0.00 sec) </code></pre> Then we can use the multiple-table <code>DELETE</code> syntax as follows: <pre class="prettyprint"><code>DELETE t2 FROM mytb t1 JOIN mytb t2 ON (t2.url_addr = t1.url_addr AND t2.url_id > t1.url_id); </code></pre> ... which will delete duplicate entries, leaving only the first url based on <code>url_id</code>: <pre class="prettyprint"><code>SELECT * FROM mytb; +--------+-------------------+ | url_id | url_addr | +--------+-------------------+ | 1 | www.google.com | | 2 | www.microsoft.com | | 3 | www.apple.com | | 5 | www.cnn.com | +--------+-------------------+ 3 rows in set (0.00 sec) </code></pre> <hr> UPDATE - Further to new comments above: If the duplicate URLs will not have the same format, you may want to apply the <code>REPLACE()</code> function to remove <code>www.</code> or <code>http://</code> parts. For example: <pre class="prettyprint"><code>DELETE t2 FROM mytb t1 JOIN mytb t2 ON (REPLACE(t2.url_addr, 'www.', '') = REPLACE(t1.url_addr, 'www.', '') AND t2.url_id > t1.url_id); </code></pre>

You may want to try the method mentioned at http://labs.creativecommons.org/2010/01/12/removing-duplicate-rows-in-mysql/. <pre class="prettyprint"><code>ALTER IGNORE TABLE your_table ADD UNIQUE INDEX `tmp_index` (URL_ADDR); </code></pre>

Remove duplicates using only a MySQL query?

Tags:

sql

mysql

I have a table with the following columns:

URL_ID    
URL_ADDR    
URL_Time

I want to remove duplicates on the URL_ADDR column using a MySQL query.

Is it possible to do such a thing without using any programming?

976

asked Aug 01 '10 21:08

Jim

3 Answers

Consider the following test case:

CREATE TABLE mytb (url_id int, url_addr varchar(100));

INSERT INTO mytb VALUES (1, 'www.google.com');
INSERT INTO mytb VALUES (2, 'www.microsoft.com');
INSERT INTO mytb VALUES (3, 'www.apple.com');
INSERT INTO mytb VALUES (4, 'www.google.com');
INSERT INTO mytb VALUES (5, 'www.cnn.com');
INSERT INTO mytb VALUES (6, 'www.apple.com');

Where our test table now contains:

SELECT * FROM mytb;
+--------+-------------------+
| url_id | url_addr          |
+--------+-------------------+
|      1 | www.google.com    |
|      2 | www.microsoft.com |
|      3 | www.apple.com     |
|      4 | www.google.com    |
|      5 | www.cnn.com       |
|      6 | www.apple.com     |
+--------+-------------------+
5 rows in set (0.00 sec)

Then we can use the multiple-table DELETE syntax as follows:

DELETE t2
FROM   mytb t1
JOIN   mytb t2 ON (t2.url_addr = t1.url_addr AND t2.url_id > t1.url_id);

... which will delete duplicate entries, leaving only the first url based on url_id:

SELECT * FROM mytb;
+--------+-------------------+
| url_id | url_addr          |
+--------+-------------------+
|      1 | www.google.com    |
|      2 | www.microsoft.com |
|      3 | www.apple.com     |
|      5 | www.cnn.com       |
+--------+-------------------+
3 rows in set (0.00 sec)

UPDATE - Further to new comments above:

If the duplicate URLs will not have the same format, you may want to apply the REPLACE() function to remove www. or http:// parts. For example:

DELETE t2
FROM   mytb t1
JOIN   mytb t2 ON (REPLACE(t2.url_addr, 'www.', '') = 
                   REPLACE(t1.url_addr, 'www.', '') AND 
                   t2.url_id > t1.url_id);

200

answered Oct 02 '22 20:10

Daniel Vassallo

You may want to try the method mentioned at http://labs.creativecommons.org/2010/01/12/removing-duplicate-rows-in-mysql/.

ALTER IGNORE TABLE your_table ADD UNIQUE INDEX `tmp_index` (URL_ADDR);

answered Oct 02 '22 18:10

Box

This will leave the ones with the highest URL_ID for a particular URL_ADDR

DELETE FROM table
WHERE URL_ID NOT IN 
    (SELECT ID FROM 
       (SELECT MAX(URL_ID) AS ID 
        FROM table 
        WHERE URL_ID IS NOT NULL
        GROUP BY URL_ADDR ) X)   /*Sounds like you would need to GROUP BY a 
                                   calculated form - e.g. using REPLACE to 
                                  strip out www see Daniel's answer*/

(The derived table 'X' is to avoid the error "You can't specify target table 'tablename' for update in FROM clause")

answered Oct 02 '22 18:10

Martin Smith

Related questions
                            
                                Multiple SQL Update Statements in single query
                            
                                MySQL "Unknown Column in On Clause" [duplicate]
                            
                                I forgot the semicolon ";" in a MySQL Terminal query. How do I exit?
                            
                                The variable name '@' has already been declared. Variable names must be unique within a query batch or stored procedure. in c#
                            
                                How would I implement a simple site search with php and mySQL?
                            
                                Finding unmatched records with SQL
                            
                                MySQL.. Return '1' if a COUNT returns anything greater than 0
                            
                                T-SQL average rounded to the closest integer
                            
                                Best way to defend against mysql injection and cross site scripting
                            
                                Is there a better way to calculate the median (not average)
                            
                                How do I delete blank rows in Mysql?
                            
                                New line in TSQL
                            
                                Error in phpmyadmin Warning in ./libraries/plugin_interface.lib.php#551
                            
                                Order by day_of_week in MySQL
                            
                                How to check if a given data exists in multiple tables (all of which has the same column)?
                            
                                How to count occurrences of a node in SQL XML?
                            
                                SQL Server 2005 - Check for Null DateTime Value
                            
                                Update Field When Not Null
                            
                                String to java.sql.Date
                            
                                #1366 - Incorrect integer value:MYsql

Donate For Us

If you love us? You can donate to us via Paypal or buy me a coffee so we can maintain and grow! Thank you!

Donate Us With