DataMadeClean
← Guides

CSV vs XLSX for data processing

The choice between CSV and XLSX is usually made by whatever exported the file, not by anyone deciding. It is worth understanding anyway, because the two formats fail in opposite directions — and most real data loss happens not inside either format but in the conversion between them.

The fundamental difference

A CSV is a text file. It contains values and separators, and nothing else — no types, no formatting, no formulas, no record of what any column is meant to be. An XLSX is a zip archive of XML that stores each cell's value and its declared type alongside it.

That single difference drives everything below. A CSV cannot tell a reader that 03051 is a postcode rather than the number 3,051, so every program that opens one has to guess. An XLSX can, and does.

CSVXLSX
Stores cell typesNoYes
Stores encodingNo — must be guessedYes — internal
Multiple sheetsNoYes
Formulas and formattingNoYes
Readable without softwareYesNo
Row limitNone inherent1,048,576
Diff-able in version controlYesNo

Neither column is the good one — they are good at opposite things.

What CSV is good at

Universality and inspectability. Every language, database and tool reads CSV without a library, the file can be opened in any text editor to see exactly what it contains, and it streams — a 5 GB CSV can be processed a line at a time by a program that never holds more than one row in memory.

It is also the only one of the two that works properly in version control. Because it is text, a diff shows which rows changed; an XLSX shows as a changed binary blob.

What XLSX is good at

Preserving intent. Because each cell carries its own type, a postcode column saved as text comes back as text, on any machine, with no import settings to get right. That single property removes the majority of the problems in this guide.

It also holds things a CSV structurally cannot: multiple sheets in one file, formulas, number formats, and column widths. If a human is going to open the file, XLSX is usually the kinder choice.

The postcode problem

This is the most common conversion casualty, and it happens in one direction only: CSV to spreadsheet.

In the CSVOpened in ExcelSaved back to CSV
0305130513051
0800800800
0412345678412345678412345678

The loss happens at column two. By the time you save, the original value is already gone.

The critical detail is the timing: the damage occurs when the file is opened, before you have edited anything, because Excel decides each column's type during the import and a column of digits looks numeric. Reformatting the column as Text afterwards does not restore the zeros, since the value is now genuinely the number 800.

An XLSX is immune to this, but only if the column was text when the workbook was written. Converting a damaged CSV to XLSX does not recover anything — it faithfully preserves the damage.

The date-serial problem

Spreadsheets store dates as a day count from a fixed origin, which is why a date column sometimes displays as a five-digit number. In an XLSX, the cell type says "this is a date" and the number is interpreted correctly. In a CSV, there is no type to say so:

Cell containsXLSX readsCSV exports
3 April 2026A dateWhatever the display format was
461153 April 202646115

The CSV export writes the displayed text, not the underlying value — so the display format becomes the data.

This makes the CSV export of a date column silently dependent on the machine that produced it. The same workbook exported on a machine set to MM/DD/YYYY and one set to DD/MM/YYYY produces two different files. If you control the export, set the column's display format to ISO (YYYY-MM-DD) before exporting, and the problem disappears.

The encoding problem

An XLSX carries its own encoding internally, so an accented character survives regardless of who opens it. A CSV does not record its encoding anywhere, which means every reader guesses — and a wrong guess produces "café" where "café" should be.

This is the clearest single advantage XLSX has. There is no CSV setting that fixes it, because the format has nowhere to record the answer. The practical mitigations are to always write UTF-8, and to import via Data → From Text/CSV where the encoding can be stated explicitly rather than assumed.

Worth knowing: a UTF-8 CSV written with a byte-order mark opens correctly in Excel by double-click, and one written without it often does not. That single invisible marker at the start of the file is the difference between a working export and a support ticket.

Choosing between them

SituationUse
Feeding a database or scriptCSV, UTF-8
Data with leading zeros or mixed typesXLSX
A person will open and read itXLSX
File is over a million rowsCSV
Tracking changes in version controlCSV
Sending to someone who will open it in ExcelCSV, UTF-8 with BOM
Sending to someone whose tools you don't knowCSV, UTF-8 — say whether a BOM is included

The recurring theme: CSV for machines, XLSX for anything with types worth preserving.

The one rule worth following regardless: minimize the number of conversions. Every hop between formats is an opportunity for a type to be re-guessed, and the damage is cumulative and silent. If a file has to reach a database, exporting once from the source system directly to the format the database wants beats a round trip through a spreadsheet.

Common questions

Is XLS the same as XLSX? No. XLS is the pre-2007 binary format, and it is worth converting away from — it caps at 65,536 rows and many current tools have dropped support for reading it.

Does saving a CSV as XLSX improve the data? No. It preserves whatever state the data is in at that moment, including damage already done when the CSV was opened. Convert at the source, not after.

Why is my XLSX so much smaller than the CSV? Because it is zip-compressed, and spreadsheet data compresses very well — repeated strings are stored once in a shared table. A tenfold difference is normal and not a sign anything is missing.

Which should I ask a supplier to send? CSV in UTF-8, with dates in ISO format and any identifier columns quoted. That combination removes almost everything in this guide, and it is a reasonable thing to ask for.

Most of the corruption above happens in the gap between formats — the moment a file is opened by something that has to guess what its columns mean.

DataMadeClean reads CSV and XLSX directly. It attempts to detect CSV encoding and writes cleaned CSV output in UTF-8. XLSX cleaning reads the first worksheet; formulas, visual formatting and other sheets are not preserved. Column detection and cleaning rules still apply. See how it handles an Excel file →