Strip Jupyter Outputs Before Git (Without Breaking Your Notebook)
Why notebooks explode in pull requests, what to strip vs keep, and when nbstripout is better than a one-off web cleaner.

Git is good at text. It is bad at notebooks that contain photographs of plots encoded as base64. If your pull request is 4,000 lines of JSON and 40 lines of Python, the review is already lost.
This guide is a workflow, not a second landing page. One-off cleanup: /ipynb-output-cleaner. Extra size: /ipynb-compressor. Daily habit: nbstripout on your machine.
What actually lands in Git
An .ipynb is JSON. Each code cell can hold outputs: stdout, stderr, HTML, and image data. Matplotlib often stores a full PNG. Plotly can store even more. ipywidgets stash state under metadata. None of that is your analysis. It is a cache of the last run.
Reviewers cannot read a 12 MB JSON diff. Hosting and clone times suffer. Secrets are the worse case: an API key printed in an output cell will sit in Git history even after you “fix” the notebook later.
The one-off path (borrowed laptop)
- Keep a private copy of the notebook with outputs if you still need the figures.
- Open
/ipynb-output-cleaner, upload the working file, clear outputs (and execution counts if you want a quiet diff). - If the file is still huge, run
/ipynb-compressoron the cleaned copy. - Open the result in
/ipynb-viewerand confirm code and Markdown are intact. - Commit that file.
The cleaner processes the upload for that request and does not keep the notebook. This is not a substitute for a git hook.
The habit path (your own repo)
Install nbstripout and enable it as a filter or pre-commit hook. Then every commit strips outputs automatically. That is the correct design for a team that lives in notebooks.
Use the web cleaner when you cannot install anything: a lab PC, a Chromebook, a file someone emailed you 20 minutes before the deadline.
What to keep, what to drop
| Keep | Drop before public Git | |------|-------------------------| | Source code | Image outputs | | Markdown narrative | Widget / Plotly blobs | | Kernel metadata | Execution counts (optional, but diffs get quieter) | | Tiny text outputs you truly need in the story | Tracebacks that contain paths or tokens |
If the notebook is the paper and the figures must ship with the repo, put rendered figures in figures/ as real files and keep the notebook clean. Do not abuse Git as object storage.
Secrets
Search outputs for tokens before you push. Clearing outputs is the fastest way to drop a key you accidentally printed. Rotate the key anyway. Assume anything that once entered Git is burned.
Related tools
- Repair JSON that will not open:
/ipynb-repair - Merge several cleaned labs:
/ipynb-merger - Decision on Word vs PDF vs Markdown after Git is clean: format guide
FAQ: notebooks and Git
For public and shared repos, yes almost always. Keep a private copy or a results branch if you still need the figures. Do not commit 20 MB of PNG blobs to main.
No. The cleaner removes outputs and optionally execution counts. The compressor also minifies JSON and can drop widget state. Start with the cleaner; compress if Git still complains about size.
It should not. Source cells and Markdown stay. Only serialized outputs and, if you choose, execution counts go away.

