Why do obvious discoveries take decades?

we don't talk about how much LLM's have changed information retrieval for us. we used to have the world's knowledge in our pockets but it was still gatekept by the amount of time it took to find what you needed in impersonalized browsing. our parents told us they'd flip through hundreds of pages of encyclopaedias to find one piece of information, that we were lucky we could just search things up on the internet - but that was just the first step into instant knowledge access. the internet was a way of finding answers for questions in minutes not hours. we are now, with LLMs, finding answers in seconds not minutes.

more than the time advantage of LLMs, interdisciplinary knowledge has become fundamentally unblocked. interlinking cognitive neuropsychiatry and geographical zoning was a massive effort of self-orchestrated, time-consuming paper-reading to even find mere correlation hypotheses. topics that had nothing to do with each other did not touch each other unless a human dedicated hours of their life to finding the links. interdisciplinary knowledge and exploration was under heavy gatekeeping because it was entirely human-dependent, linearly correlated to papers read = knowledge discovered.

LLMs have made this obsolete. a model that can read hundreds of papers in two entirely separate fields and pattern-match mutual impact gets us to the questions we should be asking in seconds. it used to take us years.

every scientist knows that the most important part of the job is to ask the right questions. the research is the execution. the question is born in decades, through the specific type of intuition that blends deep intellectual creativity, pattern-matching, unconscious personal experiences, relationships and childhood curiosities that shape a question. a question is years in the making.

intellectual curiosity and discovery takes shape through the collision of these thousands of different pieces of knowledge, usually between fields, and between researchers who don't frequent any of the same circles, departments and conferences. when an LLM can collide thousands of papers together in one prompt, we get there in seconds.

this makes me very, very excited. my company exists because we've gathered massive amounts of data and can find treatment for men's fertility in seconds with our tech. while what we build isn't based on LLMs but mathematical models, it's LLMs accelerating every possible interaction between molecular biology, systemic health, tissue histology, genomics, metabolic chemistry, that accelerated our interdisciplinary team's initial paths to explore and discover our conclusions in how we build precision medicine for men. it's exactly the collision of unlinked knowledge in the dozens of different scales of the human body that have brought us to our findings and actual optimization in men's health.

the greatest danger of progress is disciplinary isolation. ironically, the more experienced a researcher, the more stuck in their ways they tend to become, the less collision in knowledge ensues.

for example; we are in a crisis of antibiotic resistance. and I believe AI has a chance to slow it down because of the reasons above.

in 2020, a team at MIT taught a model one thing and one thing only: does this molecule stop bacteria from growing. not how, not why, not what an antibiotic is supposed to look like. then they pointed it to a library of general drugs - not antibiotics, just all unused drugs - that had already failed. it pointed at a diabetes drug and said: this one here stops bacteria from growing. a diabetes drug. they renamed it halicin. it kills Acinetobacter baumannii, C. diff, tuberculosis.

E. coli developed resistance to a classic antibiotic, ciprofloxacin, in just three days. after thirty days against halicin, nothing. literally defenseless. turns out this diabetes drug collapses the proton gradient across the bacterial membrane, and kills it by hijacking its ATP machinery.

the molecule already existed and it was sitting in an unused diabetes drugs library. we didn't lack the compound, we lacked the question, because the compound had a disciplinary location: metabolic disease, not infectious disease. and a knowledge's address decides which journal reads you, which conference, which grant committee, which researchers. a compound filed under diabetes is read by people trained to study diabetes, not antibiotics. the model wasn't a human. it had no field to defend and no precedent to be loyal to. it had never been told what an antibiotic looks like, only what an antibiotic does, so it collided a diabetes drug with a gram-negative pathogen in a few hours of compute.

LLMs and deep learning are both going to find patterns everywhere: in language, in semantics, in drugs, in molecular biology, in population data. everywhere. and that's what's going to solve the world's biggest problems. i believe it. we have the solutions already - they're just waiting to be discovered. and we've built the greatest discoverers of all the tech we have ever built.

this means the moat for companies is not the collision or the discovery. anyone can collide, and collider companies don't have anything more than claude or chatgpt do. every researcher on earth will be able to slam molecular biology into tissue histology into population genomics in a single prompt, and the questions they come up with will all be equally correct and equally credulous. the defensible thing is execution in science: a protocol, a cohort, a readout at a fixed interval, and a null result you're willing to publish. discovery becomes commodity. data becomes the product.

obvious discoveries won't take decades anymore. the key now is to not make real-world validation take decades either.

Next
Next

Where the Sun is death