When you attach a spreadsheet, ChatGPT writes and runs Python on it in a sandbox. You are getting real computation, not a language model guessing at arithmetic.
1. Upload
Attach the file to a new chat. CSVs and spreadsheets are capped at roughly 50 MB, and results are much better when the file is tidy: clear column names in row one, one record per row, no merged header cells.
2. Ask what is in it first
Profile this CSV before analyzing anything:
- row and column counts
- data type and % missing for each column
- min/max/median for numeric columns
- the 5 most common values for each text column
Then list anything that looks wrong.
Gives you a factual map of the file, so you find the broken date column before you build a chart on it.
3. Clean it
Clean this data:
- trim whitespace and title-case the customer names
- parse every date to YYYY-MM-DD, flag rows you can't parse
- drop exact duplicate rows and tell me how many you dropped
Save the result as cleaned.csv and give me the download link.
Asking it to report what it dropped is the important part — that line turns a silent transformation into something you can check.
4. Chart it
Chart monthly revenue by region as a line chart,
one line per region, and tell me which region grew fastest.
ChatGPT can merge datasets on shared keys, run real statistics (t-tests, ANOVA), and produce bar, line, pie, and scatter charts — some as interactive charts you can hover, others as static images.
What it does well
Merging files, spotting missing or malformed values, reshaping columns, one-off statistics, and turning 40,000 rows into a paragraph you can put in an email.
What it can’t do
- Reach your live data. It only sees the file you uploaded. No database, no API.
- Remember the sandbox forever. The environment resets; a file uploaded hours ago may be gone.
- Be trusted on definitions. It will happily average a column that shouldn’t be averaged, or read a “12/07” date as December. It does not know your business rules unless you state them.
- Handle confidential data safely. Anything you upload leaves your machine. For customer records or anything under NDA, use a local model instead.
5. Always verify
Ask for the total row count before and after every step, and spot-check five rows by hand against the original file. Data work fails quietly — a filter that dropped 3,000 rows looks exactly like one that dropped none.
Next: clean up messy spreadsheet data or reduce hallucinations in AI answers.