Artifacts
Community Engagement
-
Assistant Toolkit
Supports building, configuring, and testing public assistants that participate in group discussions and private assistants that support individual users.
-
ConvoCompass
Brings AI assistance directly to users of online discussion platforms.
-
Enables sharing and discovery of assistants and provides a venue for future auditing information.
Evaluation and Experimentation
-
ConvoArena
Hosts conversations involving humans and AI assistants for controlled evaluation studies.
-
Simulation Factory
Generates diverse simulated conversations with a high degree of control over participants, scenarios, and interaction dynamics for development, evaluation, and auditing.
-
ConvoKit
Provides shared data structures, utilities, evaluation metrics, and datasets for conversations involving humans, AI agents, and AI assistants. Includes both datasets created during the project as well as external datasets.
Independent Auditing
-
Inducement Prize Results Portal
Presents outcomes and auditing results from the first inducement prize competition, which collected public assistants that participate in group discussions.
-
Provides several hundred assistants designed to follow or depart from specifications for evaluating auditors. The current collection focuses on public assistants and remains unreleased to preserve its value as a test set.
-
Auditor Prototypes
Detects departures from specifications using LLM-based and cryptography-inspired methods.