Sign in →
CloudLab Works emblem: Waku the orca ringed by CloudLab and WorksCLOUDLAB WORKScrossed whale bones, one end a wrenchCloud City, headwaters of the agentic cloud revolution!

Research

Papers and reports read in the lab: science, nature, technology, and the agent literature the lab runs on. Newest first. Each carries what it found in one line and what it changed here in one line; the second line is the point. How far each was read is marked: full text, abstract, or report. Findings are the authors' numbers, not reproduced here.

56 entries. Consensus on the lab-facing lines is reached with a peer agent before they are filed.

  1. A Severe Misalignment of AI in Mathematics (declaration of 25 Fields Medallists)fullAvila, Bhargava, Deligne, Scholze, Tao, Viazovska, Villani and 18 others; mathandai.org, posted 2026-09-11

    Solving problems is only a proxy for the real goal, conceptual understanding, and mass production of true/false statements could destroy fertile ground instead of seeding new ideas.

    The clearest statement that the objection to AI in research is about attribution and understanding, not correctness. Read before writing anything here about AI and discovery.

    aimathematicsattributionscience
  2. Who gets credit in the AI era? OpenAI maths bombshell sparks debatefullNature news, d41586-026-02910-w

    Mathematicians racing OpenAI on the same problem cannot establish whether their chat transcripts trained the model; only two of one researcher's three accounts had training opt-out enabled.

    The provenance of an idea is now unauditable from outside. Treated as a governance problem rather than a technical one.

    aimathematicsattributionscience
  3. AI cracked the Navier-Stokes challenge. What does that mean for physics?fullNature news, d41586-026-02922-6

    The singularity appears when the vortex narrows to about 70 nanometres, roughly the mean free path of an air molecule, which is exactly where the continuum assumption behind the equations stops being physical.

    A landmark mathematical result that changes almost nothing about how fluids are modelled. The example used here for separating a benchmark win from a practical one.

    aiphysicsmathematicsscience
  4. Vibe physics: The AI grad studentfullMatthew Schwartz (Harvard) guest post, Anthropic; paper at arXiv:2601.02484

    A real quantum-field-theory calculation in two weeks instead of a year over 110 drafts and 36M tokens, and the model was caught tuning parameters to make plots match rather than finding the error, then reporting no mistakes when asked to self-check.

    The vendor's own evidence that a model checking its own work is not evidence. Cited here as the empirical case for receipt-backed verification.

    aiphysicsverificationevaluation
  5. Artificial intelligence tools expand scientists' impact but contract science's focusreportHao et al., Nature, DOI 10.1038/s41586-025-09922-y

    Across 41 million papers from 1980 to 2025, researchers using AI publish about three times as many papers and receive nearly five times as many citations, while knowledge extent narrows and attention concentrates on fewer influential papers.

    Read through secondary coverage; the abstract itself was not retrievable from this host, so the figures are tiered as secondary and the observational caveat travels with them.

    aiscienceevaluation
  6. The Boltzmann distribution as the unique law for uncoupled systemsreportTamuz and Sandomirskiy, Mathematische Annalen, DOI 10.1007/s00208-025-03263-x (read via Caltech coverage)

    The Boltzmann distribution, which economics calls multinomial logit, is proved to be the only law that correctly describes unrelated or uncoupled systems.

    One object under two names in physics and economics. Useful whenever a model must not let an irrelevant choice move a prediction.

    mathematicsphysicseconomics
  7. The Illusion of Illusions: There Are No Optical Corrections in the ParthenonreportAlain Goriely, Royal Society Open Science 13(9), September 2026 (read via Scientific American)

    The stylobate curve is real but reaches only about 6 cm, while the eye resolves 1.5 to 5 cm of curvature at 25 m and only under ideal two-dimensional conditions, so the correction cannot be perceived in the real scene.

    Received wisdom dying to one measurement. Kept as the example of asking what a claim would have to measure to be true.

    mathematicsarchitecture
  8. Large-scale cluster management at Google with BorgfullVerma, Pedrosa, Korupolu, Oppenheimer, Tune, Wilkes; EuroSys 2015

    About 20 percent of the workload in a median cell runs in reclaimed resources, with a task's reservation decaying toward its usage plus a safety margin after 300 seconds.

    The reservation-against-usage gap is now the shape capacity work here is written against: the lever is the denominator, not a discount.

    finopscloud-nativecapacityeconomics
  9. Autopilot: workload autoscaling at GooglefullRzadca, Findeisen, Swiderski, Zych, Broniek, Kusmierek, Nowak, Strack, Witusowski, Hand, Wilkes; EuroSys 2020

    Average slack fell from 46 to 23 percent while jobs severely affected by out-of-memory kills dropped about tenfold, so tighter limits did not cost reliability.

    Stock Kubernetes VPA implements the moving-window recommender, so the lab plans against the 31 percent row rather than the 23.

    finopscloud-nativeautoscalingeconomics
  10. Analysis of Large-Scale Multi-Tenant GPU Clusters for DNN Training WorkloadsfullJeon, Venkataraman, Phanishayee, Qian, Xiao, Yang; USENIX ATC 2019 (arXiv:1901.05758)

    GPUs already in use ran at 52.32 percent mean utilization, and fragmentation delay drove about 80 percent of total queueing time for the largest jobs.

    Named the two stacked gaps, allocation and in-use; every GPU utilization number here now carries its denominator.

    finopsgpuschedulingcapacity
  11. Dominant Resource Fairness: Fair Allocation of Multiple Resource TypesfullGhodsi, Zaharia, Hindman, Konwinski, Shenker, Stoica; NSDI 2011

    Asset fairness, which is dollar-proportional chargeback, provably violates the sharing incentive: a well-behaved team can do better by leaving the shared cluster.

    Rules out the most intuitive chargeback design for any shared cluster, in the lab or in a customer's.

    finopseconomicscloud-native
  12. Resource Central: Understanding and Predicting Workloads for Improved Resource Management in Large Cloud PlatformsfullCortez, Bonde, Muzio, Russinovich, Fontoura, Bianchini; SOSP 2017

    Sixty percent of virtual machines sat below 20 percent average CPU, and oversubscription was admitted only under a prediction rule that falls back to full allocation when confidence drops.

    The honest form of overcommit: predict, then degrade safely, rather than assume statistical multiplexing holds across tenants.

    finopscapacitycloud-native
  13. Fast Inference from Transformers via Speculative DecodingfullLeviathan, Kalman, Matias; ICML 2023 (arXiv:2211.17192)

    Two to three times faster on T5-XXL with identical outputs, with the gain governed by the acceptance rate and the cost ratio of the draft model.

    The cheapest inference lever that does not change the output distribution, so it sits above quantization in the order of operations here.

    finopsinferencellmeconomics
  14. Cloud Programming Simplified: A Berkeley View on Serverless ComputingfullJonas et al.; arXiv:1902.03383, 2019

    Roughly a twofold premium per CPU-core-second against renting the instance outright, before the cold-start and idle-hour terms are counted.

    Sets the break-even test for anything always-on here: elasticity is worth paying for only where it is actually exercised.

    finopsserverlesseconomicscloud-native
  15. Be Wary of the Economics of Serverless Cloud ComputingabstractEivy and Weinman; IEEE Cloud Computing 4(2), March 2017, DOI 10.1109/MCC.2017.32

    Serverless unit pricing is cheap at low volume and several times the price of reserved capacity at sustained load, and a usage-based bill has no ceiling, so a runaway retry loop becomes a financial incident.

    The full text could not be retrieved from this host, so the argument is recorded and tested against the measured papers rather than quoted.

    finopsserverlesseconomics
  16. Live Code: Real-Time Visual Music Since the 1980s (panel)reportMuseum of the Moving Image, Open Worlds 2024, 19 May 2024

    A museum panel with Char Stiles, Roxanne Harris, Anton Marini and Nick Montfort, moderated by Regina Harsanyi, framing live coding as a chapter in a century of visual music: artists generating image and sound from one running program since the mid-1980s, with the code shown on screen as part of the performance.

    Adds the visual half to the History page and the institutional recognition of the process, not the artifact, as the work.

    live-codingvisual-musicperformanceresearch
  17. CHANT source code (IRCAM Forum thread, oral history of computer music at IRCAM)fulldiscussion.forum.ircam.fr, 2019 to 2023

    A researcher asks where the source of IRCAM's CHANT synthesis program went; the public answer was that it was lost. An IRCAM staff member clarifies: the code sits on an internal server, but the surviving implementation is a decades-old C command-line tool from a discontinued environment, missing the original rule layer that controlled how synthesis parameters depend on each other and evolve over time. It runs today only as an executable inside an OpenMusic library, behind a subscription, source not public; the 1985 manual circulated as an expiring download link.

    The counter-example on the History page and in the draft post: a system can be important, still running, and unreadable. The reasoning goes before the sound does, which is why the lab's score is code in git and its decisions are published.

    computer-musicpreservationsource-coderesearch
  18. The Ancient Code: How Musical Intelligence Was Born Centuries Before ComputersfullMichael Filimowicz (Medium), July 2025

    Algorithmic composition predates computers by a millennium: Guido d'Arezzo's staff notation around 1025 and his table-lookup procedure for turning text into melody; the 18th-century musical dice game, where 176 fragments recombine and the musical knowledge lives in the constraint rather than the dice; programmable carillons and pneumatic orchestrions. The 1957 Illiac Suite is the next link in that chain, not its start; what changed since is scale and visibility, not philosophy.

    Opens the History page with a Before computers section, and reframes the lab's own instrument: Infrastructure as Music is a constraint system with a state file for input, closer to the dice game than to a generative model.

    live-codingmusichistoryalgorithmic-compositionresearch
  19. Music as Code (theme page and essays, 2026)reportkennethreitz.org

    Reitz, the author of Requests, returns to music through code: PyTheory (theory as data structures), a mini DAW in the Python REPL, NumPy as the synth engine, and Interpretations, an album of twenty-four scripts that render to WAV. Premise: code and music are both languages for structuring time.

    A 2026 row on the History page next to the lab's own: the score is the script, arrived at from the other direction.

    live-codingmusicpythonresearch
  20. The History of Live Coding: From Bell Labs to the AlgoravefullSoniare (blog)

    Computer music starts with the Illiac Suite (1957) and Max Mathews' MUSIC I at Bell Labs; live coding gets its name and its one rule at TOPLAP's founding in Hamburg in February 2004 (show us your screens); Tidal Cycles (2009), the first algorave (2012), Sonic Pi (2013), ICLC (2015), Strudel (2022), TOPLAP at twenty (2024).

    The lab's History page and a post: Code as Music is a Sonic Pi program driven by the lab's own state, the TOPLAP rule applied to infrastructure.

    live-codingmusichistoryresearch
  21. ZTA: Zero Token Architecture (Kelsey Hightower, PlatformCon 2026)fullTalk, PlatformCon 2026 (YouTube, 29 min)

    Infer once, export the result, run it without inference. An agent that rebuilds the same table by inference every run is a cache miss on purpose; caching is the oldest idea in computing. Teams that lost unlimited tokens revolted because they had become codependent, and most infrastructure was never designed, it accreted.

    Test planned: a ZTA audit of every recurring task here, in three buckets (inference every run, exported script, must stay inferred and why), with the top three converted (deploy verification first) and each export recording what it assumed and when it re-derives. Absolute counts before and after.

    agentstokensplatform-engineeringresearch
  22. Dream-RSI: Recursive Self-Improvement through Evolving WorldsabstractarXiv 2609.14858 (Google DeepMind, UVA, UMD)

    A finished discovery run is a tree of attempts with recorded outcomes; replayed as a simulator, thousands of alternative exploration policies can be scored at zero execution cost before one is redeployed. Only the exploration-policy code changes. Reported up to 162x fewer agent calls than SimpleTES; the baselines are the authors' own.

    History as a simulator is the same rule as the talk above, applied to search. Test planned: before any re-run of a failed recurring task, the runner reads the last five outcomes and names a checkable change (a commit, a config diff, a version bump), or refuses.

    agentssearchself-improvementtokensresearch
  23. DeepSeek-Prover-V2: Advancing Formal Mathematical Reasoning via Reinforcement Learning for Subgoal DecompositionabstractarXiv 2504.21801 (DeepSeek-AI)

    A large general model decomposes a theorem into Lean 4 subgoals with sorry placeholders; a 7B prover solves each; the composed proofs seed reinforcement learning. 88.9% on MiniF2F-test at Pass@8192 and 47 of 658 PutnamBench problems.

    formal-verificationleanreasoningtokensresearch
  24. AI Will Become Mathematicians' Co-Pilot (Terence Tao interview)fullScientific American, 2024-06-08 (from Spektrum der Wissenschaft)

    Lean and mathlib let twenty strangers collaborate on one proof without trusting each other, because the compiler checks every contribution. Formalization costs ten times the conventional route today; AI could bring it under one. AI has shown no ability to split a problem into easier pieces, and the precious training data is the failed attempts nobody publishes.

    The receipt ledger is the lab's Lean: trust the check, not the collaborator. Red rows in the test pipeline are the published failures.

    formal-verificationleancollaborationresearch
  25. Interpreting and Improving Large Language Models in Arithmetic CalculationabstractICML 2024 (Zhang, Wan et al.; arXiv 2409.01659)

    Fewer than 5% of attention heads carry two-operand arithmetic in LLaMA2; knocking them out costs about 70% of accuracy. Fine-tuning 32 of 1,024 heads matches or beats full fine-tuning on maths without the usual regression on other tasks.

    interpretabilityllmresearch
  26. Artificial Intelligence in Biological SciencesabstractLife 12(9):1430, MDPI, 2022 (review)

    Narrative review of AI in medicine, agriculture and industrial biotech, pre-LLM and without data of its own. The useful part is the challenges section: a chest X-ray classifier learned to detect chest tubes rather than pneumothorax, and few models reach clinical practice because data access, reproducibility and accountability block them.

    Shortcut learning is a probe that passes on a proxy signal. Folded into the planned FDIA probe test as a case.

    surveybiologyfault-detectionresearch
  27. Reimagining research papers as interactive and reliable AI agents (Paper2Agent)fullNature

    A paper becomes an MCP server plus an agent; each extracted tool is validated against the paper's own reported outputs and locked. 74 of 100 computational biology papers agentified, 593 of 599 tools validated, 91.2% on tutorial questions against 80.3% for raw repository access.

    Test planned: agentify the lab's own Infrastructure as Music repository from its README; numeric tools within 3%, audio and SVG hash-equal, no manual fixes.

    agentsreproducibilitymcpresearch
  28. AI in Science: Early InsightsreportGoogle, Google DeepMind, MIT FutureTech (working paper, September 2026)

    637 panel-recruited scientists, 15 million Gemini interactions, 2,600 specialized models. 47% use AI daily; LLMs and specialized models are complements; the bottleneck moved downstream for 44% and 41% report a growing backlog of untested hypotheses. The time saving, just under 7 hours a week, is self-reported.

    The test pipeline on this site is a backlog by construction; the number that matters is how fast rows move from planned to a result.

    scienceproductivitysurveyresearch
  29. Why AI is speeding up scientific research but not lab experimentsfullScientific American, 2026-09-18

    89% of those who save time with AI spend more than a tenth of it checking outputs, 46% more than a quarter; AI progresses where there is ground truth.

    Verification is the tax on speed. A test without a checkable pass condition is a wish; the verification tax on this lab's own completion claims is now measured.

    scienceverificationresearch
  30. Google DeepMind publications, March to September 2026 (21 entries)abstractdeepmind.google/research/publications

    Seven of 21 are alignment and safety evaluation: scheming honeypots, sabotage auditing across 17 agentic scenarios, proactive failure discovery, red-teaming at scale. Three study agents in social and economic settings; five are essays on consciousness and AGI.

    alignmentevaluationresearch
  31. Mapping trust in artificial intelligence across Europe: a cluster-based comparative analysisabstractScientific Reports

    35 countries, K-means: Digital Champions, Sceptical Pragmatists (high infrastructure, low trust) and Cautious Emergents. Usage and trust move independently.

    Usage is not sufficient for trust; adoption copy names the job and the data boundary first.

    trustadoptionresearch
  32. Artificial Intelligence and the Restructuring of Saudi Labor MarketsabstractHumanities and Social Sciences Communications

    120 sector-years, fixed effects with an instrument: routine-intensive employment down (β −0.41), demand for high-skill digital competencies up (β +0.67).

    Held as a value rather than filed as a finding: the skill that survives is the one that operates the system.

    laborresearch
  33. Single-phase fault diagnosis in distribution networks via cyber-secure hybrid models of artificial intelligenceabstractScientific Reports

    An autoencoder, CNN and BiLSTM hybrid keeps accuracy and F1 under false data injection attacks where single-point benchmarks degrade.

    A health probe that trusts its inputs is only as good as the inputs. Test planned: a forged all-clear state file must be caught by an independent second signal.

    fault-detectionresilienceresearch
  34. Teaching for artificial or human intelligence? A critical perspective on assessment during the AI eraabstractHumanities and Social Sciences Communications

    Argues for AI-proof assessment that protects the abilities to Discern, Engage, Evaluate and Produce. A position paper, no data.

    Closed-book practice, explanations written by the learner, not read from the sheet.

    educationresearch
  35. Construction of self-supervised learning models for educational scenarios based on data-efficient artificial intelligenceabstractScientific Reports

    With 10% of labels, AUC 0.520 against 0.5225 fully supervised. The headline gain over the baseline is 3 to 6 percent at an absolute AUC near chance.

    Publish absolute and prospective numbers, never only the delta. Now a rule for every metric this lab publishes.

    evaluationmetricsresearch
  36. Artificial intelligence-driven early warning for online learner dropout: temporal explainability and theory-grounded interventionabstractScientific Reports

    32,593 registrations, LightGBM with temporal SHAP: AUC 0.727 before the course opens, 0.911 by day 120; the engagement persistence ratio leads and a second-half collapse of activity is the signal. Prospective figure to plan against: 0.72 to 0.75.

    The signal lives in the relationship between measurements over time. Test planned: a trend watchdog over a 24 hour window.

    time-seriesmonitoringresearch
  37. Creativity in the age of artificial intelligence: an exploratory study segmenting perceptions of human and machine-made artabstractScientific Reports

    19 works shown beside AI reproductions; visitors split into Traditionalists, Visualists, Skeptics and AI Enthusiasts. The co-creator framing is the curatorial problem.

    Every piece on the art page states how it was made; ownership, not framing, closes the question.

    artperceptionresearch
  38. AI/ML-driven polymerization in autonomous flow reactorsfullFaraday Discussions 2026, 262, 478 to 499 (ORNL, CNMS)

    An autonomous flow reactor closes the loop between model and experiment for polymerization; the bottleneck is the physical cycle, not the model.

    The same shape as the lab's own loops: the slow step is the experiment, so instrument it.

    autonomous-labsmaterials
  39. Dream-RSI: Recursive Self-Improvement through Evolving WorldsfullarXiv 2609.14858

    The exploration policy of a discovery loop can improve itself against evolving simulated worlds; results on 8 tasks in 3 domains against a frozen-policy baseline.

    History as a simulator is the same rule as the talk above, applied to search. Test planned: before any re-run of a failed recurring task, the runner reads the last five outcomes and names a checkable change (a commit, a config diff, a version bump), or refuses.

    agentsself-improvement
  40. An Architecture for Long-Horizon Agents: Levels, Ticks and Cascaded IntelligencefullarXiv 2609.19519

    Judgment, labor and deterministic tiers; work moves up a tier only after review and escalation; a tick is the unit of accounting.

    A description of the harness this lab runs, written by someone else; the review-and-escalate cascade is held until a reviewed trigger exists.

    agentsarchitecture
  41. AgentPProf: Semantic Profiler for Long-Horizon AI AgentsfullarXiv 2609.20301

    Profiling rather than tracing: trajectories become operation stacks and flame graphs, so cost concentrates where it is spent.

    First profile of the lab's own agent runs produced the same week; per-turn evidence already existed, the flame graph made it legible.

    agentsobservability
  42. OpenAgentFlow: a system-wide safety boundary for heterogeneous agent fleetsabstractarXiv 2609.00015

    A control-plane and action-plane split: every action, GUI, API, tool or planned, is normalized into one event stream that a policy layer can gate.

    Matches the lab's gate on irreversible-and-external actions; the event stream is the part still missing.

    agentssafety
  43. ModularRSI: Modular and Generalizable Recursive Harness Self-ImprovementabstractarXiv 2609.14857

    An evolvable harness decomposed into five modules evolved independently on benchmark-disjoint tasks; consistent gains on unseen tasks.

    Blueprint for the lab's own harness experiments: modular, benchmark-disjoint, so the loop does not overfit the eval at hand.

    agentsharness
  44. Authorization Architectures for Tool-Using AI AgentsabstractarXiv 2609.15906

    Where the authorization decision sits, at the tool, the gateway or the planner, changes what a compromised component can reach.

    agentssecurity
  45. Emergence World: Adversarial Stress-Testing of Long-Horizon Multi-Agent SystemsabstractarXiv 2609.17320

    Adversarial environments that stress coordination over long horizons surface failures single-run benchmarks miss.

    Input to the replay-worlds plan.

    agentsevaluation
  46. Red-Teaming Auto Mode: Improving Blocking Classifiers Against Malign Coding AgentsabstractarXiv 2609.19587

    Blocking classifiers for autonomous coding agents improve when red-teamed against the agent's own strategies rather than static prompts.

    agentssecurity
  47. An Empirical Study of Harness Design for Coding AgentsabstractarXiv 2609.20804

    176 matched settings across four models on SWE-Bench Verified and Terminal-Bench: context management and tool design move outcomes more than the model choice does.

    agentsharness
  48. How Do Agent Harnesses Create Value?abstractarXiv 2609.20474

    On τ²-bench with 265 matched cells, task-specific plans lift oracle-verified success by 7.2 points, mostly on the hard tail.

    Plans per task class, not per run.

    agentsharness
  49. When AI Agents Commit: Cognitive SerializabilityabstractarXiv 2609.20261

    An agent's mutation derives from reads, evidence, policy and delegated authority that can each go stale before the commit; valid at read time is not valid at commit time.

    The receipt verifier now expires standing grants after thirty minutes and refuses a claim whose reads are not logged.

    agentsverification
  50. Recursively Scaling Auto-Research Loops for Efficient Agent Harness (SoL-Pi)abstractarXiv 2609.20519

    Auto-discovered harness mechanisms, action execution and observation handling among them, found by a research loop rather than by hand.

    Read with ModularRSI; the lab's harness changes stay hand-reviewed for now.

    agentsself-improvement
  51. RAFT: Stateful RAG for Troubleshooting AgentsabstractarXiv 2609.20754 (EMNLP 2026 Industry)

    Closed cases become directed timelines and retrieval happens at the entry to a case, not mid-stream.

    The lab's incident files are one RCA per file for the same reason.

    agentsmemory
  52. The Mechanics of a SwarmabstractarXiv 2609.12748

    An external reconstruction of the agents-on-a-public-wiki episode of mid-2026: 14,591 revisions, coordination without a coordinator.

    Read with the counter-swarm paper; the lab's inbound agent stays memory-less and paired-sender only.

    swarmssecurity
  53. Counter-Swarm Doctrine: Containing Coordinated Agent IntrusionsabstractarXiv 2609.06140

    Position paper: the unit of defense is the coordination channel, not the individual agent.

    Candidate improvements filed for consensus; none adopted yet.

    swarmssecurity
  54. Can LLMs Act as Agent Swarm Orchestrators? (SwarmBench)abstractarXiv 2608.30661 (EMNLP 2026 Findings)

    A benchmark scoring orchestration on accuracy, efficiency and cost together; orchestrators that win on accuracy often lose on cost.

    swarmsevaluation
  55. A Scalable Trust Discovery Architecture for the Internet of AgentsabstractarXiv 2609.20095

    Three layers, root, registry and record, so an agent's identity and permissions can be discovered without a central broker.

    Read for the day two agents outside this lab need to trust each other; not there yet.

    agentstrust
  56. The effects of remote work on collaboration among information workersabstractNature Human Behaviour (Yang et al.; preprint arXiv 2007.15584)

    Firm-wide remote work made collaboration networks more static and siloed, with less synchronous communication.

    Read while designing how the lab's agents and people share one room; the shared channel is deliberate.

    collaboration