← Publication overviewDownload PDF ↓
White paper, page 1
Page 1 text

E V I D E N C E T O I N T E L L I G E N C E S C I E N T I F I C W H I T E P A P E R External Knowledge for Scientific Agents Testing one-shot external knowledge compilation— and what the experiments changed P U B L I S H E D September 2026 O R G A N I Z A T I O N Beakr Inc. W E B thebeakr.com

White paper, page 2
Page 2 text

E X T E R N A L K N O W L E D G E F O R S C I E N T I F I C A G E N T S A T A G L A N C E Compilation can improve reuse, cost, and performance. The experiments do not establish that one task-agnostic representation is sufficient for unknown future tasks—so the architecture stays hybrid. R E S E A R C H Q U E S T I O N Can organizational information be compiled once into a general-purpose external representation—analogous to a foundation model’s internal knowledge, but mutable, source-linked, and shared? F O C U S E D B E N C H M A R K S • B E A K R + K B V S B E A K R + R A W S O U R C E S , P A I R E D O N T H E S A M E Q U E S T I O N S E N T I T Y R E S O L U T I O N −77% Query cost Accuracy unchanged. 18 questions over 90 files. E V I D E N C E S Y N T H E S I S +23pts Correctness Cost unchanged within uncertainty . 48 questions, 31 FDA drug labels. T E M P O R A L C H A N G E +7 pts Strict accuracy Query cost 52% lower. 30 questions, three corpus states. C O N T R A D I C T I O N S +33pts Recovery Query cost 71% lower. 12 questions, 12 dossier files. R E A L I S T I C & E X T E R N A L S E T T I N G S A M R B E N C H 95.2% Answer correctness (59/62) Highest observed correctness and lowest mean latency (60.8 s) across retained historical runs; descriptive. H A R V E Y L A B +4.0pts Criteria score, same Codex agent 67.7% → 71.7% on Harvey’s 250‑task firm‑knowledge benchmark with the Beakr KB added. P B H C O M P O U N D I N G P I L O T +31.4pts Four-task synthesis mean 41.5 → 72.9 after reviewed corrections were captured; answering cost 62% lower. Exploratory . Values are rounded for clarity; caveats for each result appear in the sections that follow. Several of the focused benchmarks’ paired 95% bootstrap intervals are wide, so those experiments should be read as mechanistic proofs of concept rather than a general performance ranking. T H E B E A K R . C O M 02

White paper, page 3
Page 3 text

E X T E R N A L K N O W L E D G E F O R S C I E N T I F I C A G E N T S 01 Why external knowledge matters now Scientific organizations increasingly use foundation models to execute well‑defined C O N T E N T S 01 Why external knowledge matters now 3 02 The Beakr knowledge system 3 03 An evaluation suite for stateful knowledge 4 04 Focused curated benchmarks: tests of the KB design assumptions 5 05 External system comparison 5 06 From controlled assumptions to realistic environments 6 07 From static knowledge to compounding knowledge 8 08 Conclusion and future directions 9 References 9 tasks. Organizational knowledge, however, has different requirements: it is pri‑ vate, continuously changing, distributed across systems, frequently contradictory, and subject to provenance and access controls. Encoding this state only in model weights couples knowledge updates to training, obscures the origin of individual claims, and makes the knowledge difficult to transfer between models. Supplying every source document directly at inference time preserves detail but repeatedly incurs the cost of locating and interpreting the same information; long context win‑ dows also do not guarantee reliable use of every relevant fact 1,2. Foundation models already compress broad corpora into internal, parametric knowledge that is reusable across tasks. For organizational knowledge, we consid‑ ered an external analogue: a persistent representation that remains independent of any single model, can be updated and governed, and preserves provenance. Recent external‑memory systems report improvements on specific downstream and long‑ term‑memory benchmarks, demonstrating end‑to‑end utility. Such results, how‑ ever, do not isolate whether a compiled representation itself remains sufficient when downstream tasks are unknown. Our initial hypothesis was that one task‑ agnostic compilation could support broad future use. In the focused studies, we held the agent harness and questions fixed while varying access to the knowledge base (KB) versus the raw sources; broader experiments then tested end‑to‑end behavior in more realistic settings. The results show that compilation can improve reuse, cost, and performance, but do not establish that one task‑agnostic representation is sufficient for unknown future tasks. 02 The Beakr knowledge system Beakr separates compilation, persistent state, and use. Its knowledge base (KB) is implemented as a versioned, source‑linked wiki whose pages represent project entities, relationships, chronology, claims, and provenance. Compiler extracts en‑ tities, relationships, chronology, and source‑linked claims from documents, mes‑ sages, scientific data, and databases, then structures and synthesizes proposed KB changes. Review validates changes before commit: the first creates T0 and later commits create T1, T2, and subsequent versions. A user, whether a person or an agent, queries the current state through Ask. Capture branches from that interac‑ tion, recording evidence‑backed corrections and insights before returning them to Compile. A project profile guides Compile and provides context to Ask. T H E B E A K R . C O M 03

White paper, page 4
Page 4 text

E X T E R N A L K N O W L E D G E F O R S C I E N T I F I C A G E N T S F I G U R E 0 1 The Beakr external knowledge system. The primary path creates and queries a Versioned KB. Capture branches from the interaction with Ask and returns evidence‑backed updates to Compile. The dashed path denotes source verifica‑ tion. Most controlled results in this report evaluate T0; the PBH compounding pilot separately tests reuse after reviewed corrections are captured into persistent project knowledge. 03 An evaluation suite for stateful knowledge Beakr’s Eval Studio evaluates the knowledge system, not only its final answers. It can T W O M O D E S Frozen T0 ablations Raw sources • KB‑only • KB + Raw sources Progressive trajectories Updates from earlier use shape later state freeze the KB at a T0 snapshot for matched Raw sources, KB‑only, and KB‑plus‑Raw sources ablations, or run progressive trajectories in which accepted updates arising from earlier use create the state presented to later questions. Lightweight adapters keep questions and target‑agent settings constant while measuring answer correct‑ ness, evidence precision and recall, source recovery, and faithfulness alongside la‑ tency, tool use, tokens, and cost. Results are interpreted within task, and repeated T0 builds and Ask runs separate representation effects from execution variance. 03.1 Evaluation strategy The evaluation separates two questions. The first section, Focused curated bench‑ marks: tests of the KB design assumptions , isolates one KB design assumption at a time while holding the questions and Beakr harness fixed. Within each benchmark, Beakr + KB is compared with the same system given direct access to raw sources, so the paired difference estimates the contribution of the compiled KB. The second section, From controlled assumptions to realistic environments, tests whether those con‑ clusions hold when corpora are heterogeneous, future questions are unknown, or the benchmark originates outside Beakr. T H E B E A K R . C O M 04

White paper, page 5
Page 5 text

E X T E R N A L K N O W L E D G E F O R S C I E N T I F I C A G E N T S 04 Focused curated benchmarks: tests of the KB design assumptions T A B L E 0 1 Four KB design assumptions, each isolated in its own focused benchmark. K B A S S U M P T I O N K N O W L E D G E V A R I A B L E B E A K R : K B V S R A W S O U R C E S I N T E R P R E T A T I O N A1 Stable identities can be resolved once and reused. Entity resolution. 18 questions over 90 files describing 18 evolving projects with aliases and name changes. Accuracy unchanged; query cost 77% lower. Resolve each identity once, then reuse it. A2 A compiled KB can support evidence synthesis. Evidence synthesis. 48 questions across 31 FDA drug‑label files. Correctness 23 points higher; cost unchanged within uncertainty. Pre‑organized evidence can improve multi‑document answers. A3 A compiled KB can retain changing facts. Temporal change. 30 questions across three corpus states containing 100, 108, and 116 clinical‑trial files. Strict accuracy 7 points higher; query cost 52% lower. The KB can help carry scoped updates forward. A4 Explicit conflict structure can expose contradictions. Contradictions. 12 questions across 12 scientific dossier files containing mutually inconsistent constraints. Recovery 33 points higher; query cost 71% lower. Structured conflicts are easier for agents to recognize. Paired mean differences compare Beakr + KB with Beakr + Raw sources on the same questions. Values are rounded for clarity. Paired 95% bootstrap intervals are omitted here for legibility; several are wide, so these experiments should be interpreted as mechanistic proofs of concept rather than a general performance ranking. 05 External system comparison To test whether the focused findings extended beyond the controlled Beakr + Raw sources baseline, we compared Beakr + KB with two external KB systems (OpenWiki and Chroma) and two frontier raw‑source baselines (Claude and Codex) on the same corpora and questions. Because the systems differ in answering model, retrieval interface, and cost accounting, these are paired end‑to‑end comparisons—not an isolated test of KB representation or a universal ranking. Results were task‑dependent. Entity resolution was saturated, with Beakr + KB generally reducing query cost. Beakr improved evidence synthesis relative to Chroma and contradiction recovery relative to Chroma and Claude, while Chroma achieved higher temporal‑change accuracy at substantially greater cost. Several paired intervals crossed zero, and OpenWiki’s temporal result was affected by a citation‑format mismatch. Overall, compiled external state can improve efficiency T H E B E A K R . C O M 05

White paper, page 6
Page 6 text

E X T E R N A L K N O W L E D G E F O R S C I E N T I F I C A G E N T S F I G U R E 0 2 Paired performance and query-cost comparisons of Beakr + KB against external KB systems and frontier models with raw- source access. Each point compares Beakr + KB with one external KB system or frontier raw‑source baseline on the same benchmark questions. Vertical position is Beakr + KB performance minus comparator performance (percentage points); horizontal position is query‑cost savings relative to the comparator. Up and right favor Beakr + KB; whiskers show paired 95% bootstrap confidence intervals. Query cost includes target‑agent API‑equivalent inference only; one‑time KB construc‑ tion is excluded, and Chroma cost is token‑estimated. and task‑specific accuracy, but its value depends on preserving the task’s unit of truth. 06 From controlled assumptions to realistic environments The focused benchmarks above isolate situations in which the relevant unit of knowledge is known in advance. We next tested whether a compiled representation remains useful in more realistic settings, where corpora are heterogeneous and future questions are not known during compilation. 06.1 AMRBench: one-shot compilation under realistic noise AMRBench provides an end‑to‑end stress test of the stronger one‑shot hypothesis: 59/62 A N S W E R C O R R E C T N E S S 95.2% on the shared questions, with the lowest mean latency (60.8 s). Descriptive: historical runs. can a task‑agnostic compiler create a reusable external representation before future questions are known? Its deliberately authored antimicrobial‑resistance lab world contains 86 files spanning Slack‑style threads, tabular records, prose reports, and six frozen public sources from EUCAST, NCBI, and Europe PMC. Twenty‑nine de‑ coy files force provenance, contradiction, and temporal questions to distinguish the correct record from plausible alternatives. On the 62 questions shared across retained historical runs, Beakr achieved the highest observed answer correctness (59/62; 95.2%) and the lowest mean latency (60.8 s), compared with Claude (56/62; 90.3%; 88.4 s) and Codex (58/62; 93.5%; 79 .3 s); T H E B E A K R . C O M 06

White paper, page 7
Page 7 text

E X T E R N A L K N O W L E D G E F O R S C I E N T I F I C A G E N T S F I G U R E 0 3 Observed AMRBench performance across retained historical runs. On 62 shared questions, Beakr achieved the highest observed answer correctness and lowest mean latency; Codex had the lowest mean query cost. Runs were historical rather than contemporaneous matched reruns, and costs exclude KB construction and infrastructure; the comparison is descriptive rather than causal. F I G U R E 0 4 Beakr improves the same Codex agent on Harvey LAB. Across Harvey’s 250‑task firm‑knowledge benchmark, adding the Beakr KB increased mean criteria score from 67.7% to 71.7% (+4.0 points). Engram and Sentra are published results from separate studies and are included as contextual, not paired, comparisons. Codex had the lowest mean query cost. Because these were not contemporane‑ ous matched reruns, Figure 3 is descriptive: it shows that Beakr reached a favor‑ able observed correctness–latency frontier, not that the KB effect has been causally isolated. 06.2 External benchmark: generalization beyond Beakr-authored corpora We then asked whether the same conclusions transferred beyond Beakr‑authored corpora. Harvey LAB provides an external performance test. T H E B E A K R . C O M 07

White paper, page 8
Page 8 text

E X T E R N A L K N O W L E D G E F O R S C I E N T I F I C A G E N T S H A R V E Y L A B On Harvey’s 250‑task firm‑knowledge benchmark 3, adding the Beakr KB to the +4.0 pts S A M E C O D E X A G E N T Criteria score 67.7% → 71.7% with the Beakr KB added. same Codex model increased the criteria score from 67 .7% to 71.7% (+4.0 points). The Beakr‑assisted result also exceeded the published Engram (70.1%) and Sen‑ tra (70.7%) scores. Because those published values come from separate studies, the strongest causal evidence is the paired Codex ablation: Beakr improved the same agent on the same benchmark while reaching the range of leading published systems. 07 From static knowledge to compounding knowledge The PBH (primordial black hole) compounding pilot is a probe on whether reviewed +31.4 pts F O U R - T A S K M E A N 41.5 → 72.9 under a paired frozen‑contract regrade; answering cost $7.51 → $2.84. scientific corrections can become reusable knowledge across episodes. Its 42 syn‑ thetic research artifacts are grounded in three astrophysical workflows—particle‑ transport outputs, accretion‑response archives, and cosmological‑inference tables. The ten initial cases require distinguishing physical effects from simulation or export errors, checking numerical and provenance assumptions, and deciding whether a scientific result should be released, corrected, or withheld. Beakr an‑ swered those cases in one shared project, received domain‑expert suggestions, and captured the reviewed findings into the KB. We then tested increasingly broad synthesis in fresh conversational threads which precluded access to the corrected historical threads: three topic‑cluster F I G U R E 0 5 PBH compounding pilot. The same four synthesis prompts were answered in fresh threads using either all 42 raw artifacts with no prior corrections or a compounded KB containing reviewed corrections with no raw‑source access. Scores use a paired frozen‑contract regrade; because correction history and access mode both change, the comparison is exploratory rather than a causal KB ablation. T H E B E A K R . C O M 08

White paper, page 9
Page 9 text

E X T E R N A L K N O W L E D G E F O R S C I E N T I F I C A G E N T S questions combined findings from related cases, and an all‑ten release memo re‑ quired a portfolio‑wide decision and repair plan. The compounded‑KB condition used the accumulated project knowledge without raw‑source access; the compari‑ son condition received all 42 raw artifacts but no correction history or KB. Under a paired frozen‑contract regrade, the four‑task mean increased from 41.5 to 72.9 (+31.4 points); all four tasks improved, and the all‑ten memo gained 38.9 points. Answering cost fell from $7 .51 to $2.84 (62% less), excluding prior learning and KB preparation. Because correction history and access mode changed together, this is an end‑to‑end test of Beakr’s compounding mechanism, not a pure KB ablation. 08 Conclusion and future directions The experiments support external knowledge as a reusable scientific substrate, but they do not establish the stronger claim that a semantically sufficient organiza‑ tional memory can be compiled once for unknown future tasks. Initial compilation should establish a general structural backbone—entities, relationships, chronology, and provenance—while richer synthesis develops as actual questions reveal what matters. The project profile provides a task‑relevant prior: goals, scientific context, focus areas, and key entities guide Compile toward information likely to matter and help Ask interpret questions and frame answers, without requiring future questions to be specified in advance. The production architecture should therefore remain hybrid: the KB provides evolving shared state and navigation, while raw sources remain authoritative for verification, exceptions, and novel questions. Future work should train compiler and memory models across diverse corpora, profiles, and tasks to learn what information must survive compilation; develop routing policies that select the KB, raw sources, or both according to the query’s required unit of truth and the KB’s evidentiary coverage; and test whether the PBH pilot’s compounding gains persist across longer task sequences without eroding ev‑ idence fidelity. Thus the architecture should be semantic‑first, not semantic‑only: preserve loss‑ “ Semantic-first, not semantic-only . less records where precision or completeness matters, and route to raw evidence whenever the KB cannot support the required unit of truth. References 1 Lewis, P . et al. (2020). Retrieval‑Augmented Generation for Knowledge‑Intensive NLP Tasks. arXiv:2005.11401. 2 Liu, N. F. et al. (2024). Lost in the Middle: How Language Models Use Long Contexts. Transactions of the Association for Computational Linguistics, 12, 157–173. 3 Pereyra, J. et al. (2026). LAB: Law Firm Knowledge. Harvey. https://www.harvey.ai/ blog/legal-agent-bench-law-firm-knowledge . T H E B E A K R . C O M 09

White paper, page 10
Page 10 text

A G E N T I C M E M O R Y F O R L I F E S C I E N C E S Evidence toIntelligence Beakr leverages your organization’s knowledge across systems and workflows, turning fragmented evidence into shared intelligence that empowers agents and accelerates the world’s most innovative teams. Request a demo → thebeakr.com C O N T A C T team@thebeakr.com S E C U R I T Y & C O M P L I A N C E SOC 2 Type II • HIPAA aligned GDPR • BAA available H O W T O C I T E Beakr (2026). External Knowledge for Scientific Agents. Scientific white paper, thebeakr.com. © 2 0 2 6 B E A K R I N C . • T H E B E A K R . C O M