jupyter生成目录
Automatically generating reports is useful in a wide range of scenarios, from regularly sharing data within a company or the public, or for personal use, such as comparing the performance of different models side by side without having to manually run a Jupyter Notebook ’n’ number of times.
自动生成报告在各种情况下都很有用,从定期在公司或公众内部共享数据,或者供个人使用,例如并排比较不同型号的性能而无需手动运行Jupyter Notebook'n'次数。
In my most recent project, I wanted to be able to train several models and then calculate a set of metrics and draw result exploration plots for each. I started by building a notebook with a menu at the top which would allow me to select one of the models I had run, and then execute all the cells in the notebook to explore the results. Although it worked, this quickly became boring to do manually, so I explored how one could programmatically run Jupyter Notebooks and export HTML versions of it, which took me to the following solution.
在我最近的项目中,我希望能够训练几个模型,然后计算一组指标并为每个模型绘制结果探索图。 我首先构建了一个笔记本,该笔记本的顶部带有一个菜单,该菜单使我可以选择自己运行的模型之一,然后执行笔记本中的所有单元以探索结果。 尽管有效,但手动操作很快变得很无聊,因此我探索了如何以编程方式运行Jupyter Notebooks并导出它HTML版本,这使我进入了以下解决方案。
For this to work, you will need a template notebook where only one or two values need to be changed before execution. For this example, I am going to use this simple notebook that explores the normal distribution for different mean and standard deviations.
为此,您将需要一个模板笔记本,在执行之前,只需要更改一个或两个值即可。 对于此示例,我将使用这个简单的笔记本来探索不同均值和标准差的正态分布。
Example notebook 示例笔记本The idea with this notebook will be to programmatically change the LOC and SCALE parameters and run the rest of the cells. This will allow me to open the HTML exports side by side and explore the results.
这款笔记本的想法是通过编程方式更改LOC和SCALE参数并运行其余单元。 这将允许我并排打开HTML导出并浏览结果。
To do this, I’ll replace the 0 and 1 values for something that’s easier to parse (PUT_LOC_HERE and PUT_SCALE_HERE) and clear cell outputs to avoid conflicts in version control:
为此,我将0和1值替换为更易于解析的内容(PUT_LOC_HERE和PUT_SCALE_HERE)并清除单元格输出,以避免版本控制中的冲突:
Template 模板Then I need to set up a function that will load this notebook, replace the placeholders with the values I want to use, run the notebook and finally export it to an HTML file. This function looks as follows:
然后,我需要设置一个函数来加载此笔记本,用我要使用的值替换占位符,运行笔记本,最后将其导出到HTML文件。 该函数如下所示:
import nbformat from nbconvert.preprocessors import ExecutePreprocessor from nbconvert.exporters import HTMLExporter def create_jupyter_report(mean: float, std: float): # Load jupyter notebook into memory with open('TEMPLATE-Report.ipynb', 'r') as f: nb = nbformat.read(f, as_version=4) # Replace placeholders # Jupyter notebooks are just a JSON file, so we can use the usual # method to find and replace values # Replace mean (LOC) nb['cells'][1]['source'] = nb['cells'][1]['source'].replace("'PUT_LOC_HERE'",str(mean)) # Replace std (SCALE) nb['cells'][1]['source'] = nb['cells'][1]['source'].replace("'PUT_SCALE_HERE'",str(std)) # Execute Notebook proc = ExecutePreprocessor(timeout=600, kernel_name='python3') proc.preprocess(nb) # Export to HTML file exporter = HTMLExporter() with open(f'Report_mean-{mean}_std-{std}.html', 'w') as f: f.write(exporter.from_notebook_node(nb)[0])Then, with this function, we could run a script like the following to generate reports for a set of parameters that we would like to explore:
然后,使用此功能,我们可以运行以下脚本,以生成我们想要探索的一组参数的报告:
from create_jupyter_report import create_jupyter_report params = [ (10,1), (5,2), (100,10), (35,4) ] for p in params: create_jupyter_report(mean=p[0], std=p[1])This generates four HTMLs which I can open side by side — all of which were created in less than a minute!
这将生成四个可以并排打开HTML,所有这些HTML在不到一分钟的时间内就创建了!
Hopefully this will help you save time if you periodically have to run Jupyter Notebooks exports!
如果您需要定期运行Jupyter Notebooks导出,希望这可以帮助您节省时间!
Applied Data Science Partners is a London based consultancy that implements end-to-end data science solutions for businesses, delivering measurable value. If you’re looking to do more with your data, please get in touch via our website. Follow us on LinkedIn for more AI and data science stories!
Applied Data Science Partners是位于伦敦的一家咨询公司,为企业实施端到端数据科学解决方案,并提供可衡量的价值。 如果您想对数据做更多的事情,请通过我们的网站与我们取得联系。 在LinkedIn上关注我们,了解更多人工智能和数据科学故事!
翻译自: https://medium.com/applied-data-science/full-stack-data-scientist-5-automating-report-generation-with-jupyter-notebooks-919e32e88d18
jupyter生成目录
相关资源:四史答题软件安装包exe