bioX Home Get an API key

EH14 X

One model that reads the language biology is written in — across the human genome and across the living world, including the part of it nobody has named yet. Access is by request.

What it does

It reads sequence, it writes sequence, and it designs against a target.

It reads. Give it a stretch of DNA and it tells you how expected every letter in it is. That is how you find the places evolution has refused to change over hundreds of millions of years, which is also where a primer will bind and where a mutation is most likely to matter.

It judges a change. Give it a position in the human genome and one letter changed, and it tells you how likely that change is to cause disease — and then what it saw: which splice signal the change makes or destroys, how far it sits from the edge of an intron, and which neighbouring letters it reacted to. Sequence a person and you get hundreds of changes nobody has ruled on; this puts them in order so the hours go to the ones that matter.

It places what has no name. Take a sample, sequence what is in it, and the usual method looks every read up in a catalogue — which returns nothing at all for anything nobody has described yet. EH14 X reads a sequence the way you read handwriting rather than looking it up, so something unnamed still comes back with the company it keeps.

It designs. Name a pathogen and it writes a detection probe: a short piece of DNA that sticks to that organism and to nothing else in the sample around it. What comes out is a candidate for a laboratory to test, and every answer says so.

It looks underwater

This is the part nobody else has.

Take a litre of seawater and sequence what is in it. Below five hundred metres, the overwhelming majority of what you find matches no named species — and every tool built on a reference database answers that with silence. It is the largest part of the living world and it is invisible to the instrument we use to look at it.

EH14 X has read millions of sequences from seawater and sediment, gathered over months from environmental samples that carry no taxonomy at all. That is not a dataset you buy. It means the model has learned from the distribution that is actually out there, rather than from the catalogue of what has already been named — and it can say something useful about a sample that every other tool returns empty.

For environmental monitoring, biodiversity surveys, ballast water, aquaculture health and port surveillance, that is the difference between a result and a blank.

Access

By request, and a person reads every application.

There is no console and nothing to configure — everything runs through the API. Tell us what you want to work on and we will get back to you. Research groups, institutes and companies are all welcome to ask; what we need to know is what you intend to do with it.

Every answer carries the release that produced it, so a result you publish today can be reproduced against the same model a year from now.