SmartSPIM Pipeline: A Scalable Cloud-Based Image Processing Pipeline for Light-sheet Microscopy Data
Light-sheet fluorescence imaging is a cutting-edge 3-dimensional imaging technique that allows the analysis of biological structures and processes in-vivo and in-vitro with exceptional image resolutions [1].In recent years, this technique has evolved to increase the microscope field of view and imaging throughput resulting in the generation of huge amounts of data in the range of hundreds of Gigabytes to Terabytes per specimen [2].Consequently, this surge in data volumes presents challenges related to storage, data organization, image processing, analysis, metadata management, and visualization.Thus far, laboratories have predominantly relied on on-premise computing resources and storage solutions for data processing.These solutions afford them complete control over data and resources, thereby facilitating superior latency in data accessibility and customized security measures.Nevertheless, this approach is fraught with limitations, including constrained flexibility of compute resources, external data accessibility, scalability, and disaster recovery capabilities [3].At the Allen Institute for Neural Dynamics, we are studying large populations of specimens and sharing the data to the scientific community as early as possible which led us to rely on cloud computing resources and cloud friendly formats (OMEZarr) to develop our image processing pipelines.We are engaged in the imaging process of whole mouse brains mounted on the SmartSPIM, a Selective Plane Illumination Microscope (SPIM), and commercial light-sheet microscope system developed by LifeCanvas Technologies.Imaging is usually conducted at a resolution of 2.0 m z step, with an x-y pixel size of 1.8 m.This system facilitates the acquisition of highresolution images through the utilization of various lasers and adjustable laser powers, enabling the generation of multiple channels for dataset acquisition.In this work, we introduce the SmartSPIM pipeline, an end-to-end automated cloud light-sheet image processing pipeline which uses Code Ocean and Amazon Web Services to efficiently handle vast quantities of data.To the best of our knowledge, it is the first automated cloud-based light-sheet image processing pipeline that has been successfully deployed for this purpose.This pipeline includes rapid data transfer from a high-performance on-premise storage system to the cloud, creation of metadata, flatfield correction, image correction to remove stripe noise, large-scale image stitching facilitated by TeraStitcher [4], image registration to the Allen Mouse Brain Common Coordinate Framework (CCF) [5], a cell localization module using Cellfinder [6], quantification of cell counts per CCF region, and a cloud-native visualization of the stitched 3D volume and cell distributions categorized by brain region with Neuroglancer.It is worth mentioning that raw data, code, and results are made fully accessible to the public as an integral component of fair, open and reproducible science.Our approach involves structuring our raw data with experiment metadata.Furthermore, the results of each processed dataset include metadata derived from both raw data and all subsequent image processing modules.This derived metadata comprehensively encompasses inputs, algorithmic parameters, utilized tools, timestamps, and software versions employed at each stage of the pipeline, facilitating future reproducibility efforts.Lastly, all our code is publicly available on Github.This pipeline is capable of concurrently processing multiple datasets.Typically, the dataset size varies between 1 and 2 Terabytes, contingent upon the number of channels and nature of the experiment.For instance, in the case of a dataset of 1.5 Terabytes and 3 channels, the processing times are currently as follows: 3.3 hours for applying flatfield correction and removing stripe noise, 0.2 hours for image stitching while exclusively considering image registration within overlapping image stacks in a single channel, 1.7 hours for fusing the data with OMEZarr format, 0.1 hours for registering a single channel to the CCF atlas, 5 hours for whole-brain cell localization in multiple channels, and 0.2 hours for mapping identified cells to each CCF region.
Read more