Viyan

Viyan AI

Optimizing CRISPR Screens with AssayLoop

AssayLoop uses learned policies from historical CRISPR screen data to automate experimental design, replacing manual thresholding with a predictive model.

CRISPR screens are historically constrained by manual, iterative testing where researchers must decide which gene variants to prioritize for subsequent rounds. AssayLoop addresses this by using an amortized policy trained across 1,389 historical CRISPR experiments. Instead of researchers manually filtering lists, they provide the results of an initial round to the model. The system suggests a refined set of candidates for the next round. In practice, a researcher might see an initial, noisy list of gene candidates where the top hits are obscured by high p-values due to cell-line variability. While a human scientist might try to guess a cutoff threshold based on a standard p-value, AssayLoop looks at the specific pattern of fold-changes across the entire screen and identifies a subtle group of genes that, while individually weak, are consistently enriched together in patterns that matched successful historical experiments. The system produced a 5.67-fold increase in enrichment factor compared to random selection in tests, demonstrating that automated policy learning can outperform traditional heuristic filtering.

The Architecture of Adaptive Discovery

Existing design methods typically rely on rigid thresholds or p-value rankings from a single round of data. AssayLoop replaces this with a system that combines AssayFormer, a transformer architecture, with biological priors derived from LLM-based analysis of external databases like STRING or BioGRID. The LLM extracts structured interaction graphs from these databases to characterize the functional relationships between genes. These priors represent the gene's functional neighborhood as a dense vector. This information is integrated with the transformer to inform its predictions about which gene candidates are more likely to yield meaningful results. By incorporating these biological priors, the model can account for established pathway interactions that are not immediately obvious from a single, isolated screen.

Once the wet-lab results from an initial round arrive, the model processes this data as a structured input. The transformer performs self-attention over the sequence of gene-phenotype pairs, where the inputs are the current round's p-values and fold-change metrics. To represent this structural history, each experiment is decomposed into sequences where individual gene identities are mapped to categorical indices. Associated experimental metrics are normalized into numerical embeddings. The system maps these metrics into a fixed-length representation by treating the collection of gene-phenotype pairs as a sequence, allowing the transformer to attend to how specific experimental outcomes relate to the overall screen trajectory. The model learns the conditional probability of candidate success by comparing this current trajectory against the sequences stored in the historical training set. It identifies latent patterns by calculating the similarity between the current experimental trajectory and previous experiments that yielded successful enrichments. The output is a probability distribution over the remaining library candidates, providing a ranked list for the next iteration.

Method Enrichment Factor Hit Recovery Rate (5% of Library) Strategy Type
Random Selection 1.0x ~5% Static
Standalone LLMs Low Variable Prior-only
AssayFormer Moderate Moderate Learned Policy
AssayLoop 5.67x 27.7% Amortized Hybrid

Practical Scaling and Operational Constraints

AssayLoop shifts the cognitive load from manual tuning to data preparation. If you are applying this to a new screen, the performance is constrained by the diversity of the 1,389 historical experiments the model has already ingested. The model shows transfer capabilities across different phenotype categories, but performance varies based on the underlying biological pathway. When the target phenotype or the underlying signaling pathway is absent from the training set, the model's accuracy diminishes compared to its performance on well-represented biological systems. The primary limitation is the representational reach of the training data. Because the model relies on patterns observed in previous experimental outcomes, it remains an optimization tool for exploring known biological spaces rather than an engine for generating new, independent biological laws. Future applications will depend on the expansion of the historical dataset to include more diverse and novel regulatory circuits.