High Performance Computing & AI
Practical notes on Python, JupyterHub, Kubernetes and AI for science — from the San Diego Supercomputer Center.
-
Install a BOINC server on Jetstream
BOINC is the leading platform for volunteer computing. Scientists can create a project on the platform and submit computational jobs that will be executed on computers of volunteers all over the world.
-
Use the distributed file format Zarr on Jetstream Swift object storage
Use the distributed file format Zarr on Jetstream Swift object storage
-
Use the distributed file format Zarr on Jetstream Swift object storage
Zarr is a pretty new file format designed for cloud computing, see documentation and a webinar for more details. Zarr is also supported by dask, the parallel computing framework for Dask, and the Dask team implemented storage backends for Google Cloud Storage and Amazon S3.
-
Install custom Python environment on Jupyter Notebooks at NERSC
NERSC has provided a JupyterHub instance for quite some time to all NERSC users. It is currently running on a dedicated large-memory node on Cori, so now it can access also data on Cori $SCRATCH, not only /project and $HOME. See their documentation
-
ECSS Symposium about Jupyterhub deployments on XSEDE
Note: XSEDE has been replaced by ACCESS. ECSS Symposium, 19 December 2017, Web presentation to the XSEDE Extended Collaborative Support Services.
-
Deploy scalable Jupyterhub with Kubernetes on Jetstream
The best infrastructure available to deploy Jupyterhub at scale is Kubernetes. Kubernetes provides a fault-tolerant system to deploy, manage and scale containers. The Jupyter team released a recipe to deploy Jupyterhub on top of Kubernetes, Zero to Jupyterhub.
-
Store a conda environment inside a Notebook
Last August, during the Container Analysis Environments Workshop held at Urbana-Champaign, we had discussion about reproducibility in the Jupyter Notebooks. There came out the idea of storing all the details about the Python environment inside the Notebook, in the metadata.
-
How to modify Singularity images on a Supercomputer
Singularity allows to run your own OS within most Supercomputers, see my previous post about Running Ubuntu on Comet via Singularity
-
Deploy scalable Jupyterhub on Docker Swarm mode
Jupyterhub genrally requires roughly 500MB per user for light data processing and many GB for heavy data processing, therefore it is often necessary to deploy it across multiple machines to support many users.
-
Setup automated testing on a Github repository with Travis-ci
Start from the CUDA 8 image from NVIDIA: FROM nvidia/cuda:8.0-cudnn6-devel-ubuntu16.04
-
Setup automated testing on a Github repository with Travis-ci
It is good practice in software development to implement extensive testing of the codebase in order to catch quickly any bug introduced into the code when implementing new features.
-
Deployment of Jupyterhub with Globus Auth to spawn notebook on Comet in Singularity containers
Follow the instructions at to build images from the ubuntuanacondajupyterhub.def and centosanacondajupyterhub.def definition files, or use the containers I have already built on Comet:
-
How to create pull requests on Github
Pull Requests are the web-based version of sending software patches via email to code maintainers. They allow a person that has no access to a code repository to submit a code change to the repository administrator for review and 1-click merging.
-
How to create pull requests on Github
Pull Requests are the web-based version of sending software patches via email to code maintainers. In more detail:
-
Deploy Jupyterhub on a supercomputer with SSH authentication
The best way to deploy Jupyterhub with an interface to a Supercomputer is through the use of batchspawner. I have a sample deployment explained in an older blog post: