Benchling creates software for life science teams, ranging from academic labs to some of the largest pharmaceutical companies in the world. What started as a simple set of tools had evolved into a full platform, with over 200K scientists and 1,300 companies relying on it.
In 2025, we launched Benchling Bioanalytical, a product to help large enterprise customers test samples during drug development.
During drug development, companies run analytical tests to ensure the quality and efficacy of their drug product and manufacturing process. Developing those tests is itself a massive investment — they have to be designed and experimentally validated, just like the drug. Novel biological drugs make this even harder, since scientists are working with living materials that are sensitive and easily contaminated.
For companies like Merck, this process can be excruciating. They'd lost a major vaccine deal because their business processes and IT infrastructure couldn't keep pace with the speed of science — a bottleneck in connecting physical lab work to the data it produced.
“The thing that was taking the longest wasn't doing the experiments in lab — it was getting data from all the systems collected and analyzed.” — Roy Helmy, Associate VP, Merck
Scientists can define protocols flexibly and iterate on them until they're ready to publish.
Scientists can flexibly design plates which are used by instruments to automate lab work.
Scientists can execute work by hand or through lab automation, analyzing large volumes of data seamlessly.
Lab managers intake materials and make sure they're safely stored in the right place, with full traceability when they're used in experiments.
Lead scientists and quality teams can review work, trace back results, and manage project progress.
We decided to spend significant time on-site with Merck, doing lab tours, stakeholder interviews, and whiteboarding sessions.
The picture rarely came together cleanly. Each conversation surfaced a different system with a different interface. Each scientist was only familiar with their specific team's workflows and what happened beyond was an organizational black box. We read through hundreds of pages of protocols, raw data, and reports ourselves to piece it all together.
A major challenge was helping our teams back home understand and buy into what we learned. My PM counterpart and I hosted a series of onsite user workshops where we broke down each user journey, pulling in experts from each side to walk us through everything.
Behind any drug candidate sits a web of systems and teams that all have to stay in sync, and like most industries, the software had evolved to serve each user's specific needs, not the connections between them.
For most industries, that's an inconvenience. For drug development, the stakes are exceptionally high: incomplete data or poorly-documented work can lead regulators to pause a campaign worth billions.
Different tools capture and document data, but they don't talk to each other. Scientists end up manually chasing down samples, checking expiration dates, and recording work they'd done hours earlier. On average, a scientist needed to work across nine systems to access the information they need.
Different materials and instruments live in different systems, each with their own data structure. No one could agree on what a "sample ID" was, so they existed in four different formats, relying on know-how and Python scripts to get everything consistent. Multiply this across different systems and it led to a lot of frustration and confusion.
Robotics made execution faster, but also more data than teams could handle. Given all the manual steps in handling and processing data, the tedious work rapidly piled up. Lead scientists talked about manually reviewing reports, even on nights and weekends.
Teams didn't have visibility into each other's work, which led to confusion around timelines and requests. Every team wished the other teams understood their work better. When are my samples arriving? Where did the test results go?
The underlying reason for all this bloated tech stack was that their legacy systems couldn't operate in two modes: one for flexibly experimenting and another for executing reliably at scale.
We had the basic building blocks for flexible experimentation. After all, scientists had been doing that for years in Benchling. But, there wasn't a way to express a step-by-step process nor the ability to build guardrails when you wanted to execute it flawlessly.
And, because of that, Benchling often had similar problems as Merck. Scientists would have to pass data from one part of the app to another in order to complete an end-to-end workflow.
Data needs to move seamlessly from one step to the next without anyone having to double-check it survived the hand-off. This was the core job we had to get right and it requires us to think rigorously about how our individual features fit into the big picture.
We need to consistently ask, "What does this look like in early development vs. later when it's locked down?" We need to articulate the difference and establish patterns and language that help scientists distinguish the two.
We should always care about this, but remember the stakes are higher in drug development. We're introducing a lot of new concepts and UIs, and we need to be thoughtful what they mean and how they sync together. Ultimately, consistency builds trust in our system.
We must think about how to build on shared components and extending before re-inventing the wheel. This means getting out of team siloes and actually looking for points of leverage, whether it's a backend system or a UI component.
A major technical bet was that we could build Bioanalytical on the same foundation as another product, Bioprocess.
For years, Benchling had considered building out Bioprocess to help companies design their drug production processes. However, we observed that both teams are developing step-by-step procedures through systematic experimentation. Both needed their processes to eventually be locked down and executed.
However, existing software had evolved as separate systems, optimizing for specific use cases, while creating data silos and complex tech stacks. The more time we spent with their team, the more confident we got that we could pull it off.
My approach to leading design work was to give people real ownership, and step in deliberately when they needed it, whether that meant pairing on design explorations or troubleshooting team dynamics.
I led a design kickoff to walk through our insights and conceptual flows, then worked to balance that ownership with speed. Sometimes that meant helping a designer sift through explorations to focus on one core idea. Other times, it meant genuinely listening when someone was struggling and helping them find their footing.
Underneath it all, I just wanted to be a steady presence for the team. Things got overwhelming at times, and I tried to be someone people could count on to listen without judgment and keep them moving forward.
I want to focus on plate designs to highlight the questions around information architecture, navigation, and interaction design that I worked on with our team.
I'd originally explored a few concepts with plate design as an additional step in experiment planning. I wanted to show we could re-use existing or planned components to easily design plates, including via templates, and assign them runs (i.e. the documents where scientists would execute the work).
It turned out this wouldn't work.
Plate design was being built into our existing notebook document system. It's where our current scientists work and it hooks into other systems for registering entities, keeping audit logs and so on. Fundamentally, our experiment planning flow had grown far beyond its original intended scope and it was too late to rework those assumptions before launch.
To my surprise, Merck's scientists were okay with it.
It wasn't the ideal experience, but they were already being forced to plan experiments, design plates and execute work in different contexts, so having a single clean data flow already felt like a large improvement. It still bugged me, but I could live with the trade-off.
The silver lining was that we were able to ship improvements to how plates were visualized across our platform, including for our core R&D users. It was a great example of building with leverage and included improvements to how plate visualizations and liquid transfers.
And, even though my role was mostly to whiteboard and bounce ideas, it was nice to do some hands-on exploration and see it help all our users out.
After months of work and weeks of rigorous user acceptance testing, we saw Merck go live with a public announcement and 400 users onboarded.
Usage steadily climbed after launch based on our internal dashboards, though we didn't have time or internal staffing to dive deep into the metrics or build out better instrumentation.
One major positive signal was this led to several other teams planning to adopt Benchling, potentially expanding to thousands more users over the coming years.
Also, despite shifting focus back to bioprocessing, we landed two more prospects without any salespeople dedicated to bioanalytical. And one deal had gotten far enough where we'd basically validated our solution and were planning to close it within the next few months.
Around this same time, generative AI adoption was picking up across the org, but we made a deliberate call not to try bolting anything onto this project mid-stream.
Instead, Benchling spun up a small, dedicated AI team, with one of our founders, a group of engineers, and another designer, and I contributed use cases and feedback on emerging design patterns from the sidelines.
I was personally very interested in data entry, since scientists spend an enormous amount of time on it, including in experiment planning. What if they could just describe what they wanted, or upload a spreadsheet or screenshot that didn't match our table formats at all, and let the LLM interpret it? The hard part was balancing accuracy against token efficiency. We landed on a multi-model consensus approach for a while, which worked but was expensive to run at scale.
I ended up working on an enabling piece of that puzzle: integrating with external ontologies, so the LLM could reference a standard set of definitions and values per customer, making it more accurate without needing as many models running in consensus.
By the time I left, Benchling was gearing up to launch Benchling AI, layering generative AI across the platform to help scientists with everyday work.