Technology FAQ

Computational design of small proteins

Alces Bioworks developed a computational protein-design pipeline for creating small, purpose-built proteins with defined sequences and predicted three-dimensional structures. The questions below describe the technology and the types of design problems it can address.

What are small designed proteins?

Small designed proteins are compact proteins whose amino acid sequences are created computationally to adopt a desired structure and perform a specified molecular function. Depending on the design objective, they can be engineered to recognize a particular protein surface, bind a defined region of a target, or provide another useful molecular interaction. Designs from the Alces pipeline are typically on the order of about 100 amino acids, although size can vary with the design problem.

Antibodies are large, naturally evolved immune proteins with a characteristic architecture. Small designed proteins can be much more compact and can be built from the outset for a particular molecular target or structural objective. Their sequences and structures are computationally specified rather than generated through animal immunization or conventional antibody discovery. This can make them attractive for applications where small size, structural definition, modularity, or access to a particular protein surface is important.

No. The Alces approach is not limited to modifying a single pre-existing scaffold. Modern computational methods can generate or optimize protein structures and sequences for a specific target and design objective. The resulting proteins can therefore have architectures tailored to the molecular problem rather than being constrained to one common framework.

The pipeline combines modern methods in structural biology, machine learning, sequence design, and protein modeling. A typical workflow begins with structural information about a target, identifies candidate binding surfaces or design constraints, generates candidate protein structures and sequences, and then ranks those candidates using multiple computational metrics. The pipeline can incorporate tools for structure prediction, diffusion-based protein generation, neural-network sequence design, interface analysis, and other rapidly evolving methods.

The amino acid sequence of the target is the basic starting point. Structural information is especially valuable, including experimentally determined structures or high-confidence predicted structures. Other useful information can include the desired binding region or epitope, residues that should be contacted or avoided, relevant isoforms, species differences, post-translational modifications, oligomeric state, and the biological context in which the interaction matters.

Yes. When structural information is available, computational design can be directed toward a selected surface, pocket, domain, or set of residues. This is useful when simply binding somewhere on a protein is not sufficient and the location or geometry of the interaction matters.

The approach is potentially applicable to many proteins for which useful sequence and structural information is available. Examples can include soluble proteins, domains within larger proteins, conserved protein families, and selected surfaces of membrane-associated proteins. Feasibility depends on factors such as target structure, conformational flexibility, accessibility of the desired surface, and the physical chemistry of the intended interaction.

Potentially. Because computational design does not depend on eliciting an immune response, it can in principle explore targets or epitopes that are poorly immunogenic, highly conserved, or otherwise inconvenient for conventional immunization. The practical difficulty still depends on the target structure and the molecular characteristics of the desired binding site.

Candidate proteins can be assessed using a combination of predicted structural confidence, interface geometry, buried surface area, residue-level interactions, steric compatibility, sequence properties, and predicted behavior of the protein-target complex. No single score is definitive, so the pipeline uses multiple complementary criteria to prioritize candidates for experimental evaluation.

Designed proteins that are produced experimentally can be screened using established biophysical methods. Biolayer interferometry (BLI) and surface plasmon resonance (SPR), for example, can be used to assess whether a candidate binds its intended target and to characterize properties such as binding affinity and kinetics. Other biochemical or cell-based assays can be used when the relevant question is functional rather than simply whether binding occurs.

Yes. Because the amino acid sequence is defined, a designed protein can often be combined with common peptide tags, purification handles, labeling sites, linkers, or other genetically encoded elements. The suitability of a particular modification depends on the protein structure and on whether the added element could interfere with folding or function.

Small designed proteins can be useful wherever selective molecular recognition is valuable. Potential applications include biological research, protein detection, imaging, affinity purification, assay development, molecular characterization, and as starting points for diagnostic or therapeutic research. The requirements for each application differ, so the same design is not automatically suitable for every use.

Compact proteins can offer practical advantages in some settings, including recombinant expression, molecular diffusion, access to sterically restricted surfaces, genetic fusion to other proteins, and the construction of multivalent or multifunctional formats. Small size is not inherently better in every application, but it expands the range of architectures that can be considered.

Protein design is a predictive process, and experimental testing remains essential. Modern computational methods can sharply narrow the search space and enrich for promising candidates, but protein folding, expression, solubility, stability, and binding behavior ultimately have to be confirmed experimentally. The value of an integrated design pipeline is that computational results can be used to prioritize the most promising designs for testing and to inform subsequent design cycles.

Protein design is being transformed by advances in machine learning and structure prediction. New generative models, sequence-design methods, complex-structure predictors, and scoring approaches are appearing rapidly. The Alces pipeline was built to be adaptable so that improved methods can be incorporated as the field evolves rather than being tied to a single algorithm or software package.