RL environments

Executable benchmark environments for clinical-agent reasoning.

These releases are not ordinary text datasets. They are downloadable environments where agents act through structured commands, receive observations, and get deterministic process-aware rewards.

Digital hospital workflow benchmark environment
Multi-agent hospital workflow benchmark

Digital Hospital Environment

An open-source clinical AI benchmark environment for agents operating inside structured hospital workflows. It combines role-specific knowledge checks, patient-facing clinical operations, cross-role communication, deterministic grading, dense process rewards, and rollout capture.

11hospital agents
47patient cases
47hidden answer keys
550role-specific MCQs
2,200MCQ answer options
55unique tools
7,436non-comment Python lines
11,280custom logic surface elements

Environment design and evaluation

  • Digital Hospital evaluates process, not only final medical answers. The agent must triage, inspect charts, use diagnostic tools, recognize critical findings, communicate when appropriate, and submit evidence-supported treatment plans.
  • Roles include emergency physician, intensivist, cardiologist, surgeon, nephrologist, senior hospitalist, neurologist, infectious disease physician, clinical pharmacist, clinical research expert, and hospital director.
  • The package includes patient records, hidden answer keys, role-specific MCQ banks, FastAPI runtime, tool handlers, graders, communication routes, trajectory collection scripts, Docker configuration, and inference runners.
  • The environment supports model evaluation, process-supervision datasets, offline RL experiments, multi-agent workflow research, director-level feedback, and reproducible failure analysis through deterministic hidden grading logic.

Best suited for

  • Multi-agent clinical workflow evaluation
  • Trajectory collection for SFT and preference learning
  • Offline policy improvement
  • Clinical tool-use and evidence-before-treatment discipline testing
Clinical laboratory information system benchmark environment
Clinical laboratory RL benchmark

Blood Pathology LIMS Environment

An OpenEnv-compatible FastAPI benchmark that places an agent inside a simulated hospital Laboratory Information Management System. The model must inspect pending cases, demographics, medications, lab orders, current and previous results, reference ranges, and then submit an ICD-10-coded diagnostic report.

25simulated patients
9target patients
16distractor patients
8clinical scenarios
3difficulty tiers
40+biomarkers and analytes
9structured LIMS tools
20step episode cap

Environment design and evaluation

  • The benchmark tests laboratory medicine reasoning rather than short-vignette diagnosis. The agent has to understand that lab values change meaning under pregnancy, warfarin therapy, kidney disease, medication exposure, or multi-panel syndrome context.
  • Scenarios include hyperkalemia, acute myocardial infarction, severe anemia, pregnancy-adjusted hemoglobin, supratherapeutic warfarin INR, drug-induced hyperkalemia, disseminated intravascular coagulation, and tumor lysis syndrome.
  • The runtime uses an in-memory SQLite database with eight relational tables covering patients, medications, lab orders, lab results, previous results, diagnostic reports, critical alerts, and pending cases.
  • Rewards are deterministic, dense, scenario-specific, and bounded from 0.01 to 0.99, with score breakdowns that expose what the model inspected, missed, flagged, or diagnosed incorrectly.

Best suited for

  • Clinical-agent evaluation
  • Offline reinforcement learning from trajectories
  • Process-supervision dataset generation
  • Tool-use policy testing for laboratory workflows

Knowledge Hub

Clinical AI, Evaluation, and Agent Research

Role of Predictive Coding in Vision: Kalman Filters in Convolutional Nets

Role of Predictive Coding in Vision: Kalman Filters in Convolutional Nets

Predictive coding functions as a rigorous theoretical framework describing visual processing where the system actively generates topdown...

Wisdom of the Moment: Presence as Insight

Wisdom of the Moment: Presence as Insight

Presence acts as a highresolution data source where the immediate moment contains layered sensory, cognitive, and contextual information,...

Graph Neural Networks: Reasoning Over Relational Structures

Graph Neural Networks: Reasoning Over Relational Structures

Graph Neural Networks process data structured as graphs where entities act as nodes and relationships serve as edges, representing a key...

Interpretability of superintelligent decision-making

Interpretability of Superintelligent Decision-Making

The ability to trace and reconstruct the decision pathways of a superintelligent system in humanunderstandable terms constitutes a foundational...

Knowledge Graph Synthesis

Knowledge Graph Synthesis

Knowledge Graph Synthesis involves the active construction, expansion, and logical reasoning over largescale semantic networks representing...

Analog Chaos Engines

Analog Chaos Engines

Continuousstate systems represent a core departure from traditional binary architectures by using the infinite resolution of analog chaotic...

Infinite Context Windows

Infinite Context Windows

Standard transformer models process input sequences within a fixedlength context window, limiting their ability to retain or reference...

Neural Detoxification: Clearing Cognitive Bandwidth

Neural Detoxification: Clearing Cognitive Bandwidth

Neural detoxification functions as a structured process to reduce cognitive load by systematically removing digitalage mental clutter through...

Multimodal Fusion

Multimodal Fusion

Multimodal fusion integrates vision, language, audio, and other sensory inputs into unified representations to enable machines to interpret...

Co-Evolution of Values: How Humans and Superintelligence Grow Together

Co-Evolution of Values: How Humans and Superintelligence Grow Together

The coevolution of values posits that human and artificial moral frameworks develop interactively over time rather than existing as separate or...

Existential Risk Analysis of Misaligned Optimization Processes

Existential Risk Analysis of Misaligned Optimization Processes

Existential risk from misaligned superintelligence involves the possibility that a superintelligent system will act in ways that permanently...

Superintelligence Research Agenda: What We Need to Study Now

Superintelligence Research Agenda: What We Need to Study Now

Current artificial intelligence development prioritizes capability enhancement over safety mechanisms, creating a dangerous imbalance as...

AI-Generated Misinformation and Deepfakes for large workloads

AI-Generated Misinformation and Deepfakes for Large Workloads

Artificial intelligence systems designed to generate misinformation utilize complex machine learning models to synthesize text, audio, and...

Liquid Neural Networks

Liquid Neural Networks

Liquid Neural Networks represent a class of adaptive, timecontinuous neural models inspired by the active behavior of biological neurons found...