The invented file
This is a synthetic worked example with invented values, not a customer case or a supplier's file. Results marked computed were produced with Python's standard library on invented input, using Python 3.14's csv module; Python 3.9 gave the same results. The file is a header row and three products. The first product has a comma in two fields; the second has a doubled quote and a line break inside its description.
If the matrix is wider than the box, scroll horizontally to read every column. Keyboard: focus the matrix and use Left/Right.
sku,name,description,price
A100,"Hinge, brass","Solid brass, 50 mm",4.20
A101,"Bolt ""M6""","Zinc plated
8 pcs",1.10
A102,Washer,Flat washer,0.05A comma splitter against a real CSV reader
Splitting each line at commas gives the wrong number of fields on three of the five lines and treats the second product as two rows. A CSV reader that follows the quoting rules returns four records of four fields each, the header and the three products, and reports line 4 as the last line of the third record because that record spans lines 3 and 4. Records are numbered from the header, which is record 1. Computed with Python's csv module.
If the matrix is wider than the box, scroll horizontally to read every column. Keyboard: focus the matrix and use Left/Right.
splitting each line at commas csv reader (records)
line 1: 4 fields record 1 (line 1, header): 4 fields
line 2: 6 fields <- wrong record 2 (line 2, A100): 4 fields
line 3: 3 fields <- wrong record 3 (lines 3-4, A101): 4 fields
line 4: 2 fields <- wrong record 4 (line 5, A102): 4 fields
line 5: 4 fields
record 3 values: A101 | Bolt "M6" | Zinc plated<line break>8 pcs | 1.10The same line with the wrong delimiter
A semicolon file read with a comma delimiter produces three fields that are all wrong. Read with the declared semicolon delimiter it gives the right three. A decimal comma in a price is split in two by the wrong reader. Computed.
If the matrix is wider than the box, scroll horizontally to read every column. Keyboard: focus the matrix and use Left/Right.
line: A200;Hinge, brass;4,20
comma reader -> ['A200;Hinge', ' brass;4', '20']
semicolon reader -> ['A200', 'Hinge, brass', '4,20']A row with the wrong field count, and an unclosed quote
A row with too few or too many fields should be rejected to a list with its record number, line number, original text and reason, while the other rows import. An unclosed quote is worse, because the reader cannot say where the field ends. In a computed run with the default setting, the reader returned everything from the opening quote to the end of the file as one field, so the good rows after it vanished into that field. With strict reading it raised an error reporting unexpected end of data at the last line. Neither result tells you where the fault is, but the line after the last record that was read completely is where the unreadable record starts. So the whole file is rejected, naming that line, and no product from it is stored, including the good rows after the fault.
If the matrix is wider than the box, scroll horizontally to read every column. Keyboard: focus the matrix and use Left/Right.
file A (closed quotes): file B (one unclosed quote):
line 1: sku,name,price line 1: sku,name,price
line 2: A1,Nut,0.30 line 2: A1,"Nut,0.30
line 3: A2,Bolt line 3: A2,Bolt,1.10
line 4: A3,Washer,0.05,extra line 4: A3,Washer,0.05
file A, header = record 1 with 3 fields:
record 2 (line 2): 3 fields, expected 3 -> accepted
record 3 (line 3): 2 fields, expected 3 -> reject: too few fields
record 4 (line 4): 4 fields, expected 3 -> reject: too many fields
data rows read 3 = accepted 1 + rejected 2 (header not counted)
file B:
default reader: record 2 spans lines 2-4 and has 2 fields; the second field
is the text from Nut to the end of the file
strict reader : Error 'unexpected end of data' after reading line 4
last record read completely ended on line 1, so the unreadable
record starts on line 2 -> reject the whole file; store nothingThe assertions an importer test should make
Use the same file. Assert that four records are read, the header and three products, and that each has four fields. Assert that the stored name of the second product, A101, is exactly Bolt "M6", with the two quote marks around M6 as part of the value, and that its description contains a line break. Assert that data rows read equal rows accepted plus rows rejected on the field-count file. Assert that the semicolon version is accepted only under a declared semicolon layout and rejected under a comma layout. Assert that the file with an unclosed quote is rejected as a whole, that the message names line 2, and that none of its rows is stored, including A2 and A3. These are authored expectations, not observations of any supplier's file.
Use it to specify a priced enquiry
If your importer fails one of these, the fixed job csv-delimiter-quoting-embedded-newlines-import is £195 for one importer path and one named supplier layout. Send five to ten invented or redacted rows and the importer's name first, never real price lists, credentials or code. Prices are untested proposals, and payment follows the agreed checks and your sign-off. Nothing is booked or charged by an enquiry.
Sources and limits
- RFC 4180: common format for CSV files Checked 2026-10-11.
- Fields with line breaks, double quotes or commas should be enclosed in double quotes, and a double quote inside such a field is doubled; the document describes common practice and is not an Internet standard.
- Python csv documentation Checked 2026-10-11.
- reader.line_num counts lines read, which differs from the number of records because records can span lines; the strict option raises an Error on bad CSV input.