Biotechnology

Discovering 9 novel ketoreductases from a 250k metagenomic library

The challenge

A biotech company sought to identify novel ketoreductases (KREDs) within one of their largest metagenomic enzyme libraries.

But extracting value from such complex data required more than conventional bioinformatics search tools; it demanded a predictive modeling methodology capable of building accurate models and predicting enzyme activity amidst high sequence variability.

Without an advanced computational approach, these high-potential biocatalysts would remain hidden and inaccessible for industrial use.

 

Our solution

To unlock these novel sequences, our team conducted an Enzyme Discovery campaign that combined the use of our Biomatchmaker® software together with an AI-driven EC prediction tool our team at Zymvol had recently developed.

 

Phase 1

Filtering

We initiated the process by screening a vast pool of 250,000 SDR sequences.

To refine this library, we employed two distinct methodologies:

    1. Traditional Filtering
      Conducted by the biotech company, this method narrowed the pool down to approximately 10,000 sequences.
    2. Physics-AI Filtering
      Using the AI-driven EC predictor, this approach yielded a set of roughly 7,000 sequences.

 

Phase 2

Selection and validation

To identify the top candidates for experimental testing, we applied Biomatchmaker® technology and selected 50 high-potential sequences for laboratory validation.

Out of this selection, 22 sequences were successfully expressed in the lab.

 

Results & impact

  • Novel ketoreductases found
    Experimental validation confirmed that 9 sequences out of 22 (40% hit rate) exhibited activity against at least one of the five substrates tested, demonstrating our approach’s ability to identify functional biocatalysts.
  • Exceptional sequence diversity
    The identified sequences shared an average sequence identity of only 31% with the company’s existing KRED panel. This low degree of homology highlights the ability of the selection framework to navigate untapped regions of the sequence space, successfully uncovering novel enzymes that fall outside the scope of traditional discovery methods.

Performance Highlights

9
Novel ketoreductases
Confirmed hits showed activity for at least one of the five substrates being tested.

31%
Sequence identity
Identified highly unique hits that significantly differ from currently known KRED panels.
Page background