The choice between CSV and XLSX is usually made by whatever exported the file, not by anyone deciding. It is worth understanding anyway, because the two formats fail in opposite directions — and most real data loss happens not inside either format but in the conversion between them.
The fundamental difference
A CSV is a text file. It contains values and separators, and nothing else — no types, no formatting, no formulas, no record of what any column is meant to be. An XLSX is a zip archive of XML that stores each cell's value and its declared type alongside it.
That single difference drives everything below. A CSV cannot tell a reader that 03051 is a postcode rather than the number 3,051, so every program that opens one has to guess. An XLSX can, and does.
| CSV | XLSX | |
|---|---|---|
| Stores cell types | No | Yes |
| Stores encoding | No — must be guessed | Yes — internal |
| Multiple sheets | No | Yes |
| Formulas and formatting | No | Yes |
| Readable without software | Yes | No |
| Row limit | None inherent | 1,048,576 |
| Diff-able in version control | Yes | No |
Neither column is the good one — they are good at opposite things.
What CSV is good at
Universality and inspectability. Every language, database and tool reads CSV without a library, the file can be opened in any text editor to see exactly what it contains, and it streams — a 5 GB CSV can be processed a line at a time by a program that never holds more than one row in memory.
It is also the only one of the two that works properly in version control. Because it is text, a diff shows which rows changed; an XLSX shows as a changed binary blob.
What XLSX is good at
Preserving intent. Because each cell carries its own type, a postcode column saved as text comes back as text, on any machine, with no import settings to get right. That single property removes the majority of the problems in this guide.
It also holds things a CSV structurally cannot: multiple sheets in one file, formulas, number formats, and column widths. If a human is going to open the file, XLSX is usually the kinder choice.
The postcode problem
This is the most common conversion casualty, and it happens in one direction only: CSV to spreadsheet.
| In the CSV | Opened in Excel | Saved back to CSV |
|---|---|---|
| 03051 | 3051 | 3051 |
| 0800 | 800 | 800 |
| 0412345678 | 412345678 | 412345678 |
The loss happens at column two. By the time you save, the original value is already gone.
The critical detail is the timing: the damage occurs when the file is opened, before you have edited anything, because Excel decides each column's type during the import and a column of digits looks numeric. Reformatting the column as Text afterwards does not restore the zeros, since the value is now genuinely the number 800.
An XLSX is immune to this, but only if the column was text when the workbook was written. Converting a damaged CSV to XLSX does not recover anything — it faithfully preserves the damage.
The date-serial problem
Spreadsheets store dates as a day count from a fixed origin, which is why a date column sometimes displays as a five-digit number. In an XLSX, the cell type says "this is a date" and the number is interpreted correctly. In a CSV, there is no type to say so:
| Cell contains | XLSX reads | CSV exports |
|---|---|---|
| 3 April 2026 | A date | Whatever the display format was |
| 46115 | 3 April 2026 | 46115 |
The CSV export writes the displayed text, not the underlying value — so the display format becomes the data.
This makes the CSV export of a date column silently dependent on the machine that produced it. The same workbook exported on a machine set to MM/DD/YYYY and one set to DD/MM/YYYY produces two different files. If you control the export, set the column's display format to ISO (YYYY-MM-DD) before exporting, and the problem disappears.
The encoding problem
An XLSX carries its own encoding internally, so an accented character survives regardless of who opens it. A CSV does not record its encoding anywhere, which means every reader guesses — and a wrong guess produces "café" where "café" should be.
This is the clearest single advantage XLSX has. There is no CSV setting that fixes it, because the format has nowhere to record the answer. The practical mitigations are to always write UTF-8, and to import via Data → From Text/CSV where the encoding can be stated explicitly rather than assumed.
Worth knowing: a UTF-8 CSV written with a byte-order mark opens correctly in Excel by double-click, and one written without it often does not. That single invisible marker at the start of the file is the difference between a working export and a support ticket.
Choosing between them
| Situation | Use |
|---|---|
| Feeding a database or script | CSV, UTF-8 |
| Data with leading zeros or mixed types | XLSX |
| A person will open and read it | XLSX |
| File is over a million rows | CSV |
| Tracking changes in version control | CSV |
| Sending to someone who will open it in Excel | CSV, UTF-8 with BOM |
| Sending to someone whose tools you don't know | CSV, UTF-8 — say whether a BOM is included |
The recurring theme: CSV for machines, XLSX for anything with types worth preserving.
The one rule worth following regardless: minimize the number of conversions. Every hop between formats is an opportunity for a type to be re-guessed, and the damage is cumulative and silent. If a file has to reach a database, exporting once from the source system directly to the format the database wants beats a round trip through a spreadsheet.
Common questions
Is XLS the same as XLSX? No. XLS is the pre-2007 binary format, and it is worth converting away from — it caps at 65,536 rows and many current tools have dropped support for reading it.
Does saving a CSV as XLSX improve the data? No. It preserves whatever state the data is in at that moment, including damage already done when the CSV was opened. Convert at the source, not after.
Why is my XLSX so much smaller than the CSV? Because it is zip-compressed, and spreadsheet data compresses very well — repeated strings are stored once in a shared table. A tenfold difference is normal and not a sign anything is missing.
Which should I ask a supplier to send? CSV in UTF-8, with dates in ISO format and any identifier columns quoted. That combination removes almost everything in this guide, and it is a reasonable thing to ask for.
Most of the corruption above happens in the gap between formats — the moment a file is opened by something that has to guess what its columns mean.
DataMadeClean reads CSV and XLSX directly. It attempts to detect CSV encoding and writes cleaned CSV output in UTF-8. XLSX cleaning reads the first worksheet; formulas, visual formatting and other sheets are not preserved. Column detection and cleaning rules still apply. See how it handles an Excel file →