Imagine a web form with a set of check boxes (any or all of them can be selected). I chose to save them in a comma separated list of values stored in one column of the database table. Now, I know that the correct solution would be to create a second table and properly normalize the database. It was quicker to implement the easy solution, and I wanted to have a proof-of-concept of that application quickly and without having to spend too much time on it. I thought the saved time and simpler code was worth it in my situation, is this a defensible design choice, or should I have normalized it from the start? Some more context, this is a small internal application that essentially replaces an Excel file that was stored on a shared folder. I'm also asking because I'm thinking about cleaning up the program and make it more maintainable. There are some things in there I'm not entirely happy with, one of them is the topic of this question.

In addition to violating First Normal Form because of the repeating group of values stored in a single column, comma-separated lists have a lot of other more practical problems: <ul> <li>Can’t ensure that each value is the right data type: no way to prevent 1,2,3,banana,5 </li> <li>Can’t use foreign key constraints to link values to a lookup table; no way to enforce referential integrity.</li> <li>Can’t enforce uniqueness: no way to prevent 1,2,3,3,3,5 </li> <li>Can’t delete a value from the list without fetching the whole list.</li> <li>Can't store a list longer than what fits in the string column.</li> <li>Hard to search for all entities with a given value in the list; you have to use an inefficient table-scan. May have to resort to regular expressions, for example in MySQL: <code>idlist REGEXP '[[:<:]]2[[:>:]]'</code> or in MySQL 8.0: <code>idlist REGEXP '\\b2\\b'</code> </li> <li>Hard to count elements in the list, or do other aggregate queries.</li> <li>Hard to join the values to the lookup table they reference.</li> <li>Hard to fetch the list in sorted order.</li> <li>Hard to choose a separator that is guaranteed not to appear in the values</li> </ul> To solve these problems, you have to write tons of application code, reinventing functionality that the RDBMS already provides much more efficiently. Comma-separated lists are wrong enough that I made this the first chapter in my book: SQL Antipatterns: Avoiding the Pitfalls of Database Programming. There are times when you need to employ denormalization, but as @OMG Ponies mentions, these are exception cases. Any non-relational “optimization” benefits one type of query at the expense of other uses of the data, so be sure you know which of your queries need to be treated so specially that they deserve denormalization.

Is storing a delimited list in a database column really that bad?

Tags:

database

database-design

database-normalization

Imagine a web form with a set of check boxes (any or all of them can be selected). I chose to save them in a comma separated list of values stored in one column of the database table.

Now, I know that the correct solution would be to create a second table and properly normalize the database. It was quicker to implement the easy solution, and I wanted to have a proof-of-concept of that application quickly and without having to spend too much time on it.

I thought the saved time and simpler code was worth it in my situation, is this a defensible design choice, or should I have normalized it from the start?

Some more context, this is a small internal application that essentially replaces an Excel file that was stored on a shared folder. I'm also asking because I'm thinking about cleaning up the program and make it more maintainable. There are some things in there I'm not entirely happy with, one of them is the topic of this question.

946

asked Sep 06 '10 18:09

Mad Scientist

2 Answers

In addition to violating First Normal Form because of the repeating group of values stored in a single column, comma-separated lists have a lot of other more practical problems:

Can’t ensure that each value is the right data type: no way to prevent 1,2,3,banana,5
Can’t use foreign key constraints to link values to a lookup table; no way to enforce referential integrity.
Can’t enforce uniqueness: no way to prevent 1,2,3,3,3,5
Can’t delete a value from the list without fetching the whole list.
Can't store a list longer than what fits in the string column.
Hard to search for all entities with a given value in the list; you have to use an inefficient table-scan. May have to resort to regular expressions, for example in MySQL:
idlist REGEXP '[[:<:]]2[[:>:]]' or in MySQL 8.0: idlist REGEXP '\\b2\\b'
Hard to count elements in the list, or do other aggregate queries.
Hard to join the values to the lookup table they reference.
Hard to fetch the list in sorted order.
Hard to choose a separator that is guaranteed not to appear in the values

To solve these problems, you have to write tons of application code, reinventing functionality that the RDBMS already provides much more efficiently.

Comma-separated lists are wrong enough that I made this the first chapter in my book: SQL Antipatterns: Avoiding the Pitfalls of Database Programming.

There are times when you need to employ denormalization, but as @OMG Ponies mentions, these are exception cases. Any non-relational “optimization” benefits one type of query at the expense of other uses of the data, so be sure you know which of your queries need to be treated so specially that they deserve denormalization.

answered Oct 16 '22 10:10

Bill Karwin

"One reason was laziness".

This rings alarm bells. The only reason you should do something like this is that you know how to do it "the right way" but you have come to the conclusion that there is a tangible reason not to do it that way.

Having said this: if the data you are choosing to store this way is data that you will never need to query by, then there may be a case for storing it in the way you have chosen.

(Some users would dispute the statement in my previous paragraph, saying that "you can never know what requirements will be added in the future". These users are either misguided or stating a religious conviction. Sometimes it is advantageous to work to the requirements you have before you.)

answered Oct 16 '22 09:10

Hammerite

Related questions
                            
                                What datatype to use when storing latitude and longitude data in SQL databases? [duplicate]
                            
                                T-SQL Cast versus Convert
                            
                                Django Model() vs Model.objects.create()
                            
                                What's the difference between TRUNCATE and DELETE in SQL
                            
                                Auto Generate Database Diagram MySQL [closed]
                            
                                Rails :include vs. :joins
                            
                                Does MySQL ignore null values on unique constraints?
                            
                                PostgreSQL: Drop PostgreSQL database through command line [closed]
                            
                                Get the Last Inserted Id Using Laravel Eloquent
                            
                                What's the difference between CharField and TextField in Django?
                            
                                Compare two MySQL databases [closed]
                            
                                How to compare only Date without Time in DateTime types in Linq to SQL with Entity Framework?
                            
                                Run PostgreSQL queries from the command line
                            
                                Should each and every table have a primary key?
                            
                                When and why are database joins expensive?
                            
                                What's the best strategy for unit-testing database-driven applications?
                            
                                What database does Google use?
                            
                                How to replace a string in a SQL Server Table Column
                            
                                Best database field type for a URL
                            
                                How to remove a field completely from a MongoDB document?

Donate For Us

If you love us? You can donate to us via Paypal or buy me a coffee so we can maintain and grow! Thank you!

Donate Us With