clean

Strip superfluous metadata from notebooks

Clean notebooks before committing to remove execution counts and metadata that cause unnecessary merge conflicts. Install the hooks with nbdev-install-hooks to clean notebooks automatically. Cleaning also adds missing cell IDs, which nbformat 4.5+ requires.

Trust


source

nbdev_trust

def nbdev_trust(
    fname:str=None, # A notebook name or glob to trust
    force_all:bool=False, # Also trust notebooks that haven't changed
):

Trust notebooks matching fname.

Clean


source

clean_nb

def clean_nb(
    nb, # The notebook to clean
    clear_all:bool=False, # Remove all cell metadata and cell outputs?
    allowed_metadata_keys:list=None, # Preserve the list of keys in the main notebook metadata
    allowed_cell_metadata_keys:list=None, # Preserve the list of keys in cell level metadata
    clean_ids:bool=True, # Remove ids from plaintext reprs?
    allowed_out_metadata_keys:list=None, # Preserve the list of keys in output metadata
    repair:bool=True, # Fix structural problems first (see `repair_nb`)?
):

Clean nb from superfluous metadata

Jupyter adds a trailing \n to images in cell outputs. VS Code’s Jupyter extension does not. In this PNG example, clean_nb removes the newline to avoid editor-dependent diffs:

test_nb = read_nb('../../tests/image.ipynb')
assert test_nb.cells[0].outputs[0].data['image/png'][-1] == "\n" # Make sure it was not converted by acccident
clean_nb(test_nb)
assert test_nb.cells[0].outputs[0].data['image/png'][-1] != "\n"

The test notebook contains notebook metadata and metadata on its second cell:

test_nb = read_nb('../../tests/metadata.ipynb')

assert {'meta', 'jekyll', 'nbdev', 'my_extra_key', 'my_removed_key'} <= test_nb.metadata.keys()
assert {'meta', 'hide_input', 'my_extra_cell_key', 'nbdev', 'my_removed_cell_key'} == test_nb.cells[1].metadata.keys()

clean_nb keeps the default allowed metadata keys and removes the others:

clean_nb(test_nb)

assert {'jekyll', 'kernelspec', 'nbdev'} == test_nb.metadata.keys()
assert {'hide_input', 'nbdev'} == test_nb.cells[1].metadata.keys()

clean_nb calls repair_nb by default to fix invalid notebook structure. For example, Markdown cells must not have outputs or execution_count attributes:

_nb = dict2nb(dict(cells=[dict(cell_type='markdown', source='hi', outputs=[], execution_count=1, id='m1', metadata={})],
                   metadata=dict(kernelspec=dict(name='python3', display_name='Python 3')), nbformat=4, nbformat_minor=5))
clean_nb(_nb)
assert 'outputs' not in _nb.cells[0] and 'execution_count' not in _nb.cells[0]
validate_nb(_nb)

We can preserve some additional keys at the notebook or cell levels:

test_nb = read_nb('../../tests/metadata.ipynb')
clean_nb(test_nb, allowed_metadata_keys={'my_extra_key'}, allowed_cell_metadata_keys={'my_extra_cell_key'})

assert {'jekyll', 'kernelspec', 'nbdev', 'my_extra_key'} == test_nb.metadata.keys()
assert {'hide_input', 'nbdev', 'my_extra_cell_key'} == test_nb.cells[1].metadata.keys()

Passing clear_all=True removes everything from the cell metadata:

test_nb = read_nb('../../tests/metadata.ipynb')
clean_nb(test_nb, clear_all=True)

assert {'jekyll', 'kernelspec', 'nbdev'} == test_nb.metadata.keys()
test_eq(test_nb.cells[1].metadata, {})

clean_ids=True removes object memory addresses from text representations in outputs. These addresses can change between runs and cause unnecessary Git diffs and merge conflicts. For example:

<PIL.PngImagePlugin.PngImageFile image mode=L size=28x28 at 0x7FB4F8979690>

becomes:

<PIL.PngImagePlugin.PngImageFile image mode=L size=28x28>

clean_nb adds missing cell IDs:

test_cell = {'source': 'x=1', 'cell_type': 'code', 'metadata': {}}
_clean_cell(test_cell, False, set(), True, set())
test_cell['id']
'64e20756'

source

process_write

def process_write(
    warn_msg, proc_nb, f_in, f_out:NoneType=None, disp:bool=False
):

Directive migrations

nbdev-clean can rewrite directive comments and move directives into metadata. These migrations require explicit flags. Git hooks never run them.

Move selected directives between comments and cell metadata without changing their meaning:

c = mk_cell('#| hide\n#| eval: false\n#| export: utils\n1+1')
_to_meta(c, ['hide','eval'])
test_eq(c.metadata['nbdev'], dict(hide='true', eval='false'))
test_eq(c.source, '#| export: utils\n1+1')

_to_comments(c, ['hide','eval'])
assert 'nbdev' not in c.metadata
test_eq(c.directives, {'export':'utils', 'hide':'', 'eval':'false'})

_canon_dirs rewrites directive comments in canonical form and leaves other source unchanged. _hoist_nb_meta moves default_exp into notebook metadata. It removes cells left empty by the move, including those containing only #| hide:

c = mk_cell('#| default_exp core\n#| eval:false\n1+1')
_canon_dirs(c)
test_eq(c.source, '#| default_exp: core\n#| eval: false\n1+1')

_nb = dict2nb(dict(cells=[mk_cell('#| hide\n#| default_exp: core'), mk_cell('#| export\n1+1')],
                   metadata={}, nbformat=4, nbformat_minor=5))
_hoist_nb_meta(_nb)
test_eq(_nb.metadata['nbdev'], dict(default_exp='core'))
test_eq(len(_nb.cells), 1)
test_eq(_nb.cells[0].source, '#| export\n1+1')

source

nbdev_clean

def nbdev_clean(
    fname:str=None, # A notebook name or glob to clean
    clear_all:bool=False, # Remove all cell metadata and cell outputs?
    disp:bool=False, # Print the cleaned outputs
    stdin:bool=False, # Read notebook from input stream
    repair:<function bool_arg at 0x7f378e68d080>=True, # Fix structural problems, e.g. stray outputs on non-code cells (see `repair_nb`)?
    dirs:bool=False, # Rewrite comment directives in canonical form?
    to_meta:str=None, # Space-separated directive names to move from comments to cell metadata
    to_comments:str=None, # Space-separated directive names to move from cell metadata to comments
    nb_meta:bool=False, # Move `default_exp` into notebook metadata?
):

Clean all notebooks in fname to avoid merge conflicts

By default (fname left to None), all the notebooks in config.nbs_path are cleaned. You can opt in to fully clean the notebook by removing every bit of metadata and the cell outputs by passing clear_all=True.

To preserve extra metadata, use these settings under [tool.nbdev] in pyproject.toml:

  • allowed_metadata_keys preserves notebook metadata.
  • allowed_cell_metadata_keys preserves cell metadata.
  • allowed_out_metadata_keys preserves output metadata.

This example preserves k1 and k2 at all three levels:

[tool.nbdev]
allowed_metadata_keys = ["k1", "k2"]
allowed_cell_metadata_keys = ["k1", "k2"]
allowed_out_metadata_keys = ["k1", "k2"]

source

clean_jupyter

def clean_jupyter(
    path, model, **kwargs
):

Clean Jupyter model pre save to path

This cleans notebooks on-save to avoid unnecessary merge conflicts. The easiest way to install it for both Jupyter Notebook and Lab is by running nbdev-install-hooks. It works by implementing a pre_save_hook from Jupyter’s file save hook API.

Hooks


source

nbdev_install_hooks

def nbdev_install_hooks(
    merge:<function bool_arg at 0x7f378e68d080>=True, # Install the notebook merge driver?
    diff:<function bool_arg at 0x7f378e68d080>=True, # Install the notebook diff driver?
    globally:<function bool_arg at 0x7f378e68d080>=False, # Define the drivers in `~/.gitconfig` and the global attributes file, instead of repo files?
):

Install Jupyter and git hooks to automatically clean, trust, and fix merge conflicts in notebooks

See clean_jupyter and nbdev-merge for more about how each hook works.

Both Git drivers use the name jupyternotebook, matching nbdime. The attributes file selects this name. Git configuration determines which program runs it. You can therefore commit .gitattributes while letting collaborators use their preferred driver.

A repo installation defines the drivers in .gitconfig, included through include.path, and activates them in .gitattributes. With --globally, definitions go in ~/.gitconfig. Attributes go in core.attributesFile, which defaults to ~/.config/git/attributes.

Re-running either tool’s enable command changes its driver configuration. Git uses the definition with highest precedence. Global installations skip the repo’s post-merge trust hook.