Back to blog
22 min read
AI Breakthroughs, solved puzzles, and research problems since the public LLM era
AI Breakthroughs, solved puzzles, and research problems since the public LLM era

AI Breakthroughs, solved puzzles, and research problems since the public LLM era

[!summary] Since the public rollout of large language models in 2022–2023, AI has accelerated work across biology, chemistry, materials science, mathematics, archaeology, weather forecasting, software engineering, and creative media. The strongest examples are not just “automation,” but systems that combine pattern recognition, search, simulation, symbolic reasoning, and human expert review to make previously expensive or bottlenecked research workflows dramatically faster.

Scope and evidence standard

This note tracks major AI-enabled breakthroughs, solved or partially solved research bottlenecks, and historically important demonstrations from the public LLM era and its immediate technical lead-up.

Included breakthroughs generally meet at least one of these criteria:

  • Solved or substantially advanced a long-standing scientific or technical bottleneck.
  • Achieved expert-level or near-expert performance on a benchmark previously considered difficult for AI.
  • Enabled large-scale discovery that would have been economically or experimentally infeasible by traditional methods alone.
  • Shifted expert expectations about what machine learning systems can contribute.

[!warning] “Solved” should be read carefully. Many items below are partially solved, workflow-transforming, or benchmark breakthroughs, not complete replacements for human expertise. In science and archaeology especially, AI output still requires experimental validation, expert review, and provenance tracking.

Timeline snapshot

  • 2016–2019: AlphaGo, AlphaStar, and early protein-folding systems show deep learning + reinforcement learning can master high-dimensional strategy and scientific prediction tasks.
  • 2020–2021: AlphaFold 2 reaches near-experimental accuracy for many protein structures and changes structural biology.
  • 2022: Public LLM era begins with ChatGPT; AlphaTensor discovers new matrix multiplication algorithms; Ithaca demonstrates AI-assisted ancient text restoration and attribution.
  • 2023: GNoME predicts hundreds of thousands of stable materials; FunSearch uses LLM-guided program search for mathematical discovery; GraphCast and Pangu-Weather show AI weather forecasting can rival leading numerical systems in some settings.
  • 2024: AlphaFold 3 expands from protein structures to biomolecular interaction prediction; AlphaGeometry reaches Olympiad-level geometry performance; Vesuvius Challenge winners recover text from unopened Herculaneum scrolls.
  • 2025–2026: The trend continues toward multimodal agents, self-driving labs, AI-assisted theorem proving, scientific copilots, and increasingly automated research workflows.

Biology and chemistry

Protein folding — AlphaFold

Years: 2018–2021, with major public scientific impact continuing through 2023–2026
Organizations: DeepMind / Google DeepMind; EMBL-EBI via the AlphaFold Protein Structure Database
Key researchers and teams: John Jumper, Demis Hassabis, Richard Evans, Alexander Pritzel, Pushmeet Kohli, and the AlphaFold team

Problem

Predicting a protein’s 3D structure from its amino acid sequence was one of biology’s grand challenges for more than 50 years. Experimental methods such as X-ray crystallography, cryo-electron microscopy, and NMR spectroscopy can take months or years per protein.

Why it was hard

Protein folding involves:

  • enormous combinatorial search space
  • molecular interactions across many scales
  • sensitive 3D geometry
  • scarce experimental structures relative to known protein sequences

AI breakthrough

AlphaFold 2 achieved near-experimental accuracy for many proteins in CASP14 and was described in Nature as the first computational method able to regularly predict protein structures with atomic accuracy, even where no similar structure was known.

Historical significance

This is one of the clearest modern cases of AI breaking through a long-standing scientific bottleneck. It changed structural biology from a field constrained by experimentally solved structures into one where millions of predicted structures became available for hypothesis generation.

Technical impact

  • Structural biology at proteome scale
  • Disease and mutation interpretation
  • Enzyme engineering
  • Drug discovery support
  • Synthetic biology
  • Faster target characterization

Sources: Jumper et al., “Highly accurate protein structure prediction with AlphaFold,” Nature 2021. https://www.nature.com/articles/s41586-021-03819-2

Molecular interaction prediction — AlphaFold 3

Year: 2024
Organization: Google DeepMind / Isomorphic Labs
Key researchers and teams: Josh Abramson, Jonas Adler, Jack Dunger, Richard Evans, Tim Green, Alexander Pritzel, Olaf Ronneberger, Lindsay Willmore, John Jumper, Demis Hassabis, and collaborators

Problem

AlphaFold 2 mainly addressed protein structure prediction. Biology often depends not only on single protein structures, but on interactions among:

  • proteins
  • DNA
  • RNA
  • small-molecule ligands
  • ions
  • covalent modifications
  • glycans
  • antibody–antigen complexes

AI breakthrough

AlphaFold 3 introduced a unified, diffusion-based system for predicting joint 3D structures of broad biomolecular complexes.

Historical significance

AlphaFold 3 moved the problem from “predict the shape of a protein” toward “model molecular interaction space.” That is closer to the real operating level of cells and drug discovery.

Technical impact

Potential acceleration of:

  • drug design
  • protein–ligand modeling
  • molecular engineering
  • targeted therapeutics
  • antibody design
  • personalized medicine

Limits

AlphaFold 3 predicts static structures, not full molecular dynamics in solution. Code availability and reproducibility were also more limited than AlphaFold 2 at launch.

Sources: Abramson et al., “Accurate structure prediction of biomolecular interactions with AlphaFold 3,” Nature 2024. https://www.nature.com/articles/s41586-024-07487-w

AI-assisted drug discovery

Years: 2020–2026
Organizations: DeepMind / Isomorphic Labs, Insilico Medicine, Recursion, Exscientia, Atomwise, major pharmaceutical companies, academic labs

Problem

Traditional drug discovery is slow, expensive, and failure-prone. Most candidate compounds fail before approval.

AI breakthrough

AI systems now assist with:

  • molecule generation
  • binding affinity prediction
  • toxicity prediction
  • target identification
  • protein–ligand docking
  • compound optimization
  • trial design and patient stratification

Historical significance

This has not “solved” drug discovery, but it has changed the search economics. AI can screen and prioritize chemical space at a scale humans cannot manually inspect.

Technical impact

Potential reductions in:

  • years of lab work
  • failed compounds
  • cost per viable candidate
  • manual literature review
  • low-value wet-lab experiments

Limits

AI-generated drug candidates still require synthesis, assay validation, animal studies, clinical trials, safety review, and regulatory approval.


Materials science

Discovery of new materials — GNoME

Year: 2023
Organization: Google DeepMind
Key researchers and teams: Amil Merchant, Ekin Dogus Cubuk, Google DeepMind materials team, Materials Project collaborators

Problem

Discovering useful inorganic materials traditionally requires expensive experimentation and simulations. The chemical search space is astronomically large.

AI breakthrough

GNoME, or Graph Networks for Materials Exploration, predicted 2.2 million new crystal structures, including 380,000 stable materials considered promising for experimental synthesis.

Historical significance

DeepMind framed the result as equivalent to roughly 800 years of prior accumulated materials knowledge. The work expanded the known space of potentially stable inorganic crystals and released candidates to the Materials Project.

Technical impact

Potential applications include:

  • batteries
  • superconductors
  • solar panels
  • semiconductors
  • catalysts
  • advanced electronics
  • energy storage

Limits

Prediction is not synthesis. The most important downstream work is whether labs can make, characterize, and use the predicted materials.

Sources: Google DeepMind, “Millions of new materials discovered with deep learning,” 2023. https://deepmind.google/blog/millions-of-new-materials-discovered-with-deep-learning

Self-driving laboratories

Years: 2020–2026
Organizations: Multiple academic labs, national labs, materials science groups, chemistry automation startups

Problem

Scientific experimentation is labor-intensive, sequential, and often constrained by human trial-and-error.

AI breakthrough

AI + robotics systems can increasingly:

  • propose experiments
  • operate lab equipment
  • analyze results
  • update hypotheses
  • adapt future experiments

Historical significance

Self-driving laboratories represent a shift from AI as an offline model to AI as part of a closed-loop scientific process.

Technical impact

  • Faster experiment cycles
  • More systematic search over parameter spaces
  • Better use of expensive equipment
  • Reduced human bottlenecks in repetitive experimentation

Limits

Robotics reliability, lab safety, experimental design quality, and validation remain major constraints.


Mathematics and formal reasoning

Olympiad geometry — AlphaGeometry

Year: 2024
Organizations: Google DeepMind; New York University collaboration
Key researchers and teams: Trieu Trinh, Thang Luong, Yuhuai Wu, Quoc Le, He He, and collaborators

Problem

Olympiad geometry requires abstraction, auxiliary constructions, symbolic reasoning, proof search, and creativity. This was historically difficult for neural networks.

AI breakthrough

AlphaGeometry combined a neural language model with a symbolic deduction engine. It solved 25 of 30 International Mathematical Olympiad geometry problems from 2000–2022 under standard Olympiad time limits, near the average human gold-medalist level.

Historical significance

This was one of the clearest demonstrations that neural systems can work with symbolic logic rather than merely pattern-match text.

Technical impact

  • Stronger neuro-symbolic reasoning systems
  • Machine-verifiable proof generation
  • Better mathematical problem-solving benchmarks
  • Evidence that synthetic data can train proof-reasoning systems

Sources: Google DeepMind, “AlphaGeometry: An Olympiad-level AI system for geometry,” 2024. https://deepmind.google/blog/alphageometry-an-olympiad-level-ai-system-for-geometry

Automated theorem proving and formal math assistants

Years: 2021–2026
Organizations: OpenAI, Google DeepMind, Meta, academic Lean/Isabelle/Coq communities, mathlib contributors

Problem

Formal mathematics requires exact logical correctness. Informal proof sketches are not enough.

AI breakthrough

LLM-assisted systems increasingly help produce machine-verifiable proofs in systems such as:

  • Lean
  • Isabelle
  • Coq
  • HOL Light

Historical significance

This could eventually make proof checking and mathematical collaboration more like software engineering: versioned, testable, searchable, and partially automated.

Technical impact

  • Faster proof formalization
  • Better software verification
  • More reusable mathematical libraries
  • Reduced barrier to formal methods

Limits

LLMs still hallucinate. Formal systems catch errors, but search and translation from informal math to formal proof remain hard.

Novel mathematical constructions — FunSearch

Year: 2023 paper / 2024 Nature publication
Organization: Google DeepMind
Key researchers and teams: Bernardino Romera-Paredes, Mohammadamin Barekatain, Alexander Novikov, Matej Balog, M. Pawan Kumar, Emilien Dupont, Francisco J. R. Ruiz, Jordan S. Ellenberg, Pushmeet Kohli, Alhussein Fawzi, and collaborators

Problem

Some combinatorics and optimization problems are easy to evaluate but hard to solve. Human intuition does not search the full program space.

AI breakthrough

FunSearch paired a pretrained LLM with a systematic evaluator to search for programs. It discovered new constructions for the cap set problem and improved heuristics for online bin packing.

Historical significance

This is an important move from “AI writes plausible math text” toward “AI proposes executable objects that can be evaluated and improved.”

Technical impact

  • LLM-guided program search
  • Interpretable discovered algorithms
  • New mathematical constructions
  • Hybrid AI + evaluator loops for scientific discovery

Sources: “Mathematical discoveries from program search with large language models,” Nature 2024. https://www.nature.com/articles/s41586-023-06924-6

Algorithm discovery — AlphaTensor

Year: 2022
Organization: DeepMind
Key researchers and teams: DeepMind AlphaTensor team

Problem

Designing efficient algorithms for core computations such as matrix multiplication requires deep mathematical intuition. The search space is enormous.

AI breakthrough

AlphaTensor used reinforcement learning to discover efficient, provably correct matrix multiplication algorithms. In some cases, it improved on known algorithms, including a finite-field 4×4 case related to Strassen-style multiplication.

Historical significance

AlphaTensor showed that AI systems could discover algorithms for fundamental computational primitives, not just optimize parameters inside neural networks.

Technical impact

  • Matrix multiplication improvements
  • Hardware-aware algorithm search
  • Compiler and numerical computing implications
  • A template for automated algorithm discovery

Sources: “Discovering faster matrix multiplication algorithms with reinforcement learning,” Nature 2022. https://www.nature.com/articles/s41586-022-05172-4

Symbolic integration and equation solving

Years: ongoing, accelerated 2022–2026
Organizations: Academic labs, computer algebra communities, AI labs, tool-using LLM systems

Problem

Computer algebra systems historically depend on hand-crafted symbolic rules. They can be brittle when problems require strategy, decomposition, or translation from natural language.

AI breakthrough

Modern AI systems increasingly combine:

  • symbolic manipulation
  • theorem search
  • probabilistic reasoning
  • chain-of-thought-style decomposition
  • tool use with CAS systems

Technical impact

AI systems can assist with:

  • calculus
  • symbolic algebra
  • differential equations
  • optimization
  • physics derivations
  • translating word problems into solvable form

Limits

Formal correctness remains a hard boundary. Tool-verified computation is stronger than unsupported LLM reasoning.


Ancient languages and archaeology

Years: 2019–2026
Organizations: DeepMind / Google DeepMind; University of Oxford; University of Venice; digital humanities research groups
Key researchers and teams: Yannis Assael, Thea Sommerschield, Brendan Shillingford, and collaborators

Problem

Ancient texts are often incomplete, damaged, displaced, or linguistically ambiguous. This applies to inscriptions, cuneiform tablets, manuscripts, and fragments.

AI breakthrough

Ithaca and related systems assist with:

  • missing text restoration
  • geographic attribution
  • chronological attribution
  • handwriting recognition
  • linguistic pattern analysis
  • translation estimation

Historical significance

Ithaca showed a strong collaborative effect: historians working with the model outperformed historians alone and the model alone. This is a good example of AI as expert augmentation rather than replacement.

Technical impact

  • Faster transcription
  • More candidate restorations
  • Better search across ancient corpora
  • Scalable digital humanities workflows

Notable examples

  • Ancient Greek inscriptions
  • Akkadian cuneiform tablets
  • Dead Sea Scroll fragments
  • medieval manuscripts
  • damaged papyri

Sources: Assael et al., “Restoring and attributing ancient texts using deep neural networks,” Nature 2022. https://www.nature.com/articles/s41586-022-04448-z

Herculaneum scrolls — Vesuvius Challenge

Years: 2023–2025
Organizations: Vesuvius Challenge; University of Kentucky; international AI and imaging contributors
Key researchers and contributors: Brent Seales, Luke Farritor, Youssef Nader, Julian Schilliger, and the Vesuvius Challenge community

Problem

The Herculaneum scrolls were carbonized by the eruption of Mount Vesuvius in 79 AD. Physically unrolling them would destroy them.

AI breakthrough

Researchers used:

  • X-ray tomography
  • machine learning
  • neural imaging
  • pattern recognition
  • ink detection models

to recover text from unopened scrolls.

Historical significance

For the first time in nearly 2,000 years, scholars began reading previously inaccessible ancient writings without physically opening the scrolls.

Technical impact

  • Non-destructive reading of fragile artifacts
  • New workflow for virtual unwrapping
  • AI-assisted archaeology at prize/community scale
  • Potential recovery of entire libraries of ancient text

Sources: University of Kentucky news on Vesuvius Challenge Grand Prize, 2024. https://research.uky.edu/news/grand-prize-discovery-made-2000-year-old-herculaneum-scrolls


Language and reasoning

Emergent multi-domain reasoning in foundation models

Years: 2020–2026
Organizations: OpenAI, Anthropic, Google DeepMind, Meta, Microsoft, Mistral, xAI, academic labs, open-source communities

Problem

Researchers long expected general reasoning systems to require explicit symbolic architectures or many domain-specific systems.

AI breakthrough

Large language models demonstrated broad, imperfect, but useful reasoning across:

  • law
  • medicine
  • coding
  • education
  • psychology
  • writing
  • translation
  • planning
  • data analysis

Historical significance

This shifted the field toward foundation models, scaling laws, multimodal reasoning, and agentic systems.

Technical impact

  • General-purpose AI assistants
  • Natural-language interfaces to tools
  • AI copilots for many professions
  • Multimodal systems spanning text, image, audio, video, and code

Limits

LLMs still hallucinate, lose track over long horizons, and struggle with causality, embodiment, and reliability.

Machine translation quality

Years: Transformer era 2017 onward, accelerated by LLMs 2022–2026
Organizations: Google, Meta, Microsoft, DeepL, OpenAI, academic NLP community

Problem

Translation historically struggled with idioms, cultural nuance, context, and low-resource languages.

AI breakthrough

Transformer-based systems and LLMs dramatically improved fluency, contextual handling, and cross-lingual transfer.

Technical impact

  • Global communication
  • International collaboration
  • Multilingual software
  • Accessibility
  • Faster localization

Limits

Low-resource languages, dialects, legal/medical nuance, and cultural context still need human review.

Speech recognition and transcription

Years: 2022–2026
Organizations: OpenAI, Google, Microsoft, Meta, academic speech communities

Problem

Speech recognition struggled with accents, noise, multilingual audio, code-switching, and speaker variation.

AI breakthrough

Modern systems such as Whisper improved multilingual transcription and robustness.

Technical impact

  • Meeting transcription
  • Subtitles
  • Accessibility
  • Podcast/video indexing
  • Multilingual communication
  • Voice interfaces

Computer science and software engineering

Code generation and AI pair programming

Years: 2021–2026
Organizations: OpenAI, GitHub, Microsoft, Anthropic, Google, Cursor, Replit, Cognition, open-source coding-model communities

Problem

Software development requires large amounts of manual labor: implementation, debugging, refactoring, testing, documentation, and API learning.

AI breakthrough

LLMs can now:

  • generate code
  • refactor systems
  • debug applications
  • explain APIs
  • write tests
  • generate documentation
  • scaffold full-stack applications
  • operate coding tools as agents

Historical significance

AI coding tools are one of the most economically visible LLM-era transformations.

Technical impact

Developers increasingly use AI as:

  • pair programmers
  • architecture assistants
  • debugging collaborators
  • documentation systems
  • test writers
  • migration assistants

Limits

AI-generated code can introduce subtle bugs, security issues, dependency problems, and architecture drift. Human review and tests remain essential.

Reverse engineering and vulnerability analysis

Years: 2022–2026
Organizations: Security research labs, defense contractors, AI labs, open-source tooling communities

Problem

Analyzing large binary systems, malware, and vulnerable codebases is difficult and time-intensive.

AI breakthrough

Modern AI systems assist with:

  • malware analysis
  • binary interpretation
  • exploit detection
  • reverse engineering
  • code auditing
  • vulnerability triage

Technical impact

AI is a force multiplier for both cyber defense and offensive security research.

Limits

Security AI increases both defensive capability and misuse risk. Verification, sandboxing, and authorization boundaries matter.


Weather, climate, and physics

AI weather forecasting — GraphCast and Pangu-Weather

Years: 2023–2025
Organizations: Google DeepMind; ECMWF; Huawei; academic meteorology community
Key researchers and teams: Rémi Lam, Alvaro Sanchez-Gonzalez, Matthew Willson, Peter Battaglia, Petar Veličković, Google DeepMind GraphCast team; Huawei Pangu-Weather team

Problem

Traditional weather forecasting uses numerical weather prediction systems that solve physical equations over global grids. These are accurate but computationally expensive.

AI breakthrough

AI systems such as GraphCast and Pangu-Weather showed that learned models trained on historical atmospheric data can rival or outperform some leading numerical systems on medium-range forecasts, while running much faster.

Historical significance

Weather forecasting is a complex physical domain long dominated by physics simulation. AI weather systems demonstrated that neural models can learn useful approximations of global atmospheric dynamics.

Technical impact

  • Faster forecasts
  • Lower compute and energy cost
  • Extreme weather tracking
  • Tropical cyclone trajectory support
  • Potential democratization of forecast capability

Limits

AI weather models still depend on high-quality reanalysis/observational data and do not replace the full physical modeling, data assimilation, and operational infrastructure of meteorological agencies.

Sources: Lam et al., GraphCast paper in Science 2023; Bi et al., “Accurate medium-range global weather forecasting with 3D neural networks,” Nature 2023. https://www.nature.com/articles/s41586-023-06185-3

Plasma and fusion modeling

Years: 2020–2026
Organizations: DeepMind, EPFL, Princeton Plasma Physics Lab, national labs, fusion startups, academic plasma groups

Problem

Fusion reactors involve chaotic plasma dynamics that are difficult to predict and control.

AI breakthrough

AI systems increasingly assist with:

  • plasma stabilization
  • reactor control
  • turbulence modeling
  • magnetic confinement optimization
  • diagnostic interpretation

Technical impact

AI may accelerate practical fusion research by improving control policies, simulation speed, and experiment planning.

Limits

Fusion remains unsolved as an energy technology. AI helps with subproblems, not the full engineering and economics challenge.


Scientific research automation

AI as a scientific collaborator

Years: 2022–2026
Organizations: AI labs, universities, national labs, pharmaceutical companies, materials labs

Problem

Modern science produces more papers, datasets, methods, and code than humans can efficiently process.

AI breakthrough

Researchers now use AI for:

  • literature review
  • hypothesis generation
  • experiment planning
  • data interpretation
  • scientific writing
  • simulation orchestration
  • code generation
  • grant and manuscript drafting

Historical significance

AI is becoming a collaborative layer across the scientific method: reading, ideating, coding, testing, interpreting, and communicating.

Technical impact

  • Faster review cycles
  • Broader interdisciplinary synthesis
  • More automated data analysis
  • More reproducible computational workflows when paired with tests and provenance

Limits

Scientific claims still need evidence. AI can accelerate bad hypotheses as well as good ones if not grounded in data.

Autonomous research agents

Years: 2023–2026
Organizations: OpenAI, Anthropic, Google DeepMind, academic agent labs, open-source AI agent projects

Problem

Scientific workflows are fragmented: search, planning, coding, experimentation, evaluation, and reporting often happen in separate tools.

AI breakthrough

Agentic systems increasingly chain together:

  • search
  • planning
  • coding
  • experimentation
  • evaluation
  • reporting
  • memory and retrieval

Historical significance

This is early evidence for partially autonomous research loops, though not yet fully autonomous science.

Technical impact

  • Continuous experiment iteration
  • Automated benchmark runs
  • Self-updating literature monitors
  • Research copilots that can operate tools

Limits

Reliable long-horizon planning, error recovery, experimental grounding, and safety remain unsolved.


Games, strategy, and planning

Go — AlphaGo and successors

Years: 2016–2017
Organization: DeepMind
Key researchers and teams: David Silver, Demis Hassabis, AlphaGo / AlphaZero teams

Problem

Go was long considered too complex for brute force because of its enormous search space and strategic depth.

AI breakthrough

AlphaGo defeated elite human Go champions. AlphaZero later learned Go, chess, and shogi through self-play.

Historical significance

This was a pre-LLM-era milestone that made the later LLM-era breakthroughs more plausible: deep learning plus search could master domains requiring intuition and planning.

Technical impact

  • Reinforcement learning credibility
  • Self-play systems
  • Search-guided neural decision making
  • Inspiration for later scientific discovery systems

StarCraft II — AlphaStar

Years: 2019
Organization: DeepMind
Key researchers and teams: AlphaStar team

Problem

Real-time strategy games require incomplete information, planning, multitasking, adaptation, and long-horizon strategy.

AI breakthrough

AlphaStar achieved grandmaster-level play in StarCraft II.

Historical significance

StarCraft showed AI progress in dynamic, partially observable environments beyond board games.

Technical impact

  • Multi-agent RL
  • Real-time decision making
  • Long-horizon strategy benchmarks

Creative generation

High-quality image generation

Years: 2021–2026
Organizations: OpenAI, Stability AI, Midjourney, Adobe, Google, Meta, Runway, open-source diffusion communities

Problem

Generating realistic, controllable images from text was historically poor and unreliable.

AI breakthrough

Diffusion models and multimodal transformers dramatically improved:

  • photorealism
  • style transfer
  • composition
  • editing
  • artistic consistency
  • concept visualization

Technical impact

  • Design
  • advertising
  • entertainment
  • concept art
  • storyboarding
  • visual prototyping
  • education

Limits

Copyright, provenance, artist compensation, deepfakes, and visual misinformation remain unresolved social and legal issues.

Music and audio generation

Years: 2022–2026
Organizations: OpenAI, Google, Meta, Stability AI, Suno, Udio, ElevenLabs, Adobe, open-source audio communities

Problem

Music generation requires temporal structure, style consistency, audio quality, and controllability.

AI breakthrough

AI systems can now:

  • generate music
  • synthesize speech
  • clone voices
  • create sound effects
  • separate audio stems
  • clean noisy recordings
  • generate podcast/video narration

Technical impact

  • Media production
  • accessibility
  • localization
  • game development
  • rapid prototyping
  • voice interfaces

Limits

Voice cloning and music generation raise consent, impersonation, copyright, and attribution issues.


Things AI still has not solved

Despite rapid progress, AI has not fully solved:

  • general intelligence
  • reliable long-term planning
  • robust causal reasoning
  • hallucination prevention
  • embodied reasoning
  • common-sense grounding
  • fully autonomous science
  • full drug discovery automation
  • full mathematical creativity and proof reliability
  • true AGI

Many systems remain:

  • probabilistic
  • brittle
  • data-dependent
  • error-prone
  • hard to audit
  • heavily human-supervised

Historically important AI breakthroughs of the public LLM era

If historians ranked the most transformative AI breakthroughs around and after the public LLM era, a likely shortlist would include:

  1. AlphaFold protein structure prediction — long-standing biology bottleneck dramatically advanced.
  2. Foundation models and emergent multi-domain reasoning — general-purpose AI interface becomes public and economically visible.
  3. AI-assisted software engineering — code generation and agentic development change developer workflows.
  4. AlphaFold 3 and AI-driven molecular modeling — moves from structures toward interaction prediction.
  5. GNoME and AI materials discovery — expands candidate materials by hundreds of thousands.
  6. AlphaGeometry and formal reasoning systems — neural + symbolic systems approach elite proof problem solving.
  7. FunSearch and AlphaTensor — AI contributes executable mathematical/algorithmic discoveries.
  8. AI weather forecasting — learned models compete with major physical simulation systems in some settings.
  9. Vesuvius Challenge and ancient text restoration — AI unlocks inaccessible historical records.
  10. Generative image, video, music, and speech systems — AI becomes a major creative production layer.

Final observation

The defining trend of the post-2023 AI era is not simply automation. It is the emergence of systems capable of:

  • abstraction
  • synthesis
  • pattern discovery
  • cross-domain reasoning
  • scientific assistance
  • iterative research support
  • tool use
  • human-AI collaboration

Many researchers now treat AI as a collaborative layer across nearly every scientific and technical discipline. The most robust pattern is AI plus verification: AI proposes, searches, ranks, drafts, or predicts; humans, experiments, formal systems, or evaluators validate.


Source list

Related vault notes

  • [[ai-research-sources]]