Hacker Newsnew | past | comments | ask | show | jobs | submit | bio_lover's commentslogin

We recently launched the Latch Biodevelopment framework. Its still early days but we’d love to share it with you:

https://latchbio.substack.com/p/bio-dev-fw

Latch is a platform to build, iterate & scale biological workflows in the cloud, starting with genomics.

Why did we start Latch? We're at the tipping point of engineering the code of life around us to combat cancer, solve climate change & extend our lifespan. The rate limiting step in our ability to do this are the tools that scientists have at their disposal.

From Elliot H, "Genomics is projected to require up to 110 petabytes (PB) of storage a day within the next decade—for reference, if you were storing all of that data on 1TB hard drives, you’d need 110,000 of them per day. This would make genomics one of the largest data generating endeavors on the planet, topping other contenders such as Youtube (3-5 PB/day), and Astronomy (3 PB/day). Genomes are not the only “-ome” being measured at breakneck speed, as multi-omic studies are now generating enormous catalogs of different cellular measurements. High-resolution microscopes routinely generate terabytes (1012 bytes) of rich spatial data. In the modern life sciences, quality analytical software is utterly essential for making sense of this enormous amount of new data."

But why can't we just use existing BI tools & frameworks? Why do we need biospecific development frameworks?

A quick tour through history sketches an answer: WebDev & MLops are two great examples of the benefits of specialized development frameworks.

WebDev frameworks empowered software engineers to ship more responsive websites faster. Making a website 20 years ago meant writing raw HTML and CSS, managing the DOM state manually, & tedious reprogramming of dynamic components. Producing and hosting a website used to require a whole team of engineers. Today, developer frameworks like React, Vue, and NextJS turned all of that tediousness into managed infrastructure. Now developers focus on the website's design, architecture, creativeness, & copy.

MLOps empowered engineers to use ML in new and exciting ways. 10 years ago, training & deploying a ML model meant manually coding layers & optimization functions, rolling your hyperparameter tuning, scaling your models to the cloud, & creating observability tools to monitor distribution drifts in production. Today, developer frameworks like scikit-learn, PyTorch Lightning, HuggingFace, Tecton, & Ray turned much of that complexity into managed infrastructure. Developers can now focus on developing better architectures & prototyping new applications.

Biodevelopers deal with similar dynamics as the early adopters of WebDev & ML — lots of time wasted managing generic infrastructure.

Existing development frameworks outside biology (e.g., HuggingFace, React, NextJS) are an inspiration but don't solve the pain points of dealing with big biological data. Biology has unique file types with unique visualizations such as FASTQ, FASTA, BAM, H5AD, PDB, and GTF. "Sample sheets" encoded in CSVs are used to run a batch of workflows. Biodevelopers deal with a combination of bash scripts, C++, R, Python, and Groovy, while trying to parallelize each to hundreds of cores. The concept of automation requires interacting with physical machines from Illumina, Thermo Fisher, Oxford Nanopore, & other bio-specific providers. Data provenance & traceability are more important than in most other industries because of application processes to the FDA. And the visualization tools needed downstream are unique to biology (e.g., Seurat, IGV, ScanPy).

For those of you interested in learning more, two great articles by Elliot H:

1. The need for software in biology: https://centuryofbio.substack.com/p/whats-different-part-fou...

2. Current challenges in maintaining high quality software in biology: https://newscience.org/how-software-in-the-life-sciences-act...


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: