Contribution Prerequisites
Target Lifecycle Stage
Labs
FINOS Member Organization Name
Red Hat
Name
AI Evaluation Framework
Slug
ai-evaluation-framework
Business Problem & Solution
Business Problem
- General-purpose AI benchmarks often fall short for the unique requirements of the finance sector.
- Financial tasks rarely yield a single "correct" answer and require precise alignment with financial ontologies and regulatory standards.
- Current focus on model-level evaluations do not fully capture the complexity of real-world AI systems, which combine AI and non-AI components.
- Enterprises increasingly face a "Day 2" Reality Gap when moving AI agents into production, struggling with information overload, runaway costs, and a lack of governance and trust.
- There is a significant risk of "safety-washing," where general capability improvements in models are misinterpreted as advancements in actual system safety.
Solution
- Establish a common, open, and transparent evaluation suite for Generative AI and agentic applications in financial services, by utilizing a "Taxonomy-First" approach, mapping specific financial use cases (e.g., credit risk analysis) directly to material risks, and subsequently to measurable metrics.
- The framework bridges the trust gap by transitioning from traditional MLOps to AgentOps, enabling continuous, outcome-oriented evaluation and behavioral auditing for complex AI workflows. It acts as the critical measurement layer within the FINOS Governance-as-Code pipeline, allowing organizations to continuously evaluate their models, agents, and controls against FINOS AI Governance Framework (AIGF) risks.
Mission Statement
To establish a common, open, and transparent evaluation suite for Generative AI in financial services that bridges the AI trust gap by transitioning from model-level scoring to continuous, system-level AgentOps evaluations, mapping real-world financial use cases directly to material risks, regulatory requirements, and measurable metrics.
Strategic Objectives
- System-Level Evaluation: Move beyond generic model leaderboards to rigorously evaluate end-to-end workflows, multi-agent systems, and retrieval-augmented generation (RAG) operating in complex financial contexts.
- Taxonomy-First Framework: Systematically align financial use cases to operational and compliance risks (such as hallucinations, bias, and non-determinism) to define clear, quantitative industry benchmarks for trustworthy AI.
- Governance-as-Code Integration: Function as the open measurement layer in the FINOS Governance-as-Code pipeline, offering continuous auditability, metacognitive guardrails, and runtime evaluation across the system lifecycle.
- Mutualized Innovation: Accelerate enterprise-grade AI adoption across the financial sector through open synthetic datasets, repeatable test cases, and shared reference architectures.
Existing Materials
Core Framework Repo: LINK
Reference Implementation: LINK (The FinSight AI Agent, a multi-agent, metacognitive system designed for earnings call analysis).
Leaderboard Repo: LINK (A framework for evaluating LLM performance across diverse financial tasks).
OSFF Toronto Training Workshop: LINK
Additional Information (Optional)
This will be a multi-repo project with ai-eval-framework holding the framework itself and the various use-case implementation contributed in their own repository as software.
Maintainer Team
Compliance & Requirements Agreement
Administrative Consent & Transfer Authorization
Socialization Acknowledgement
Contribution Prerequisites
Target Lifecycle Stage
Labs
FINOS Member Organization Name
Red Hat
Name
AI Evaluation Framework
Slug
ai-evaluation-framework
Business Problem & Solution
Business Problem
Solution
Mission Statement
To establish a common, open, and transparent evaluation suite for Generative AI in financial services that bridges the AI trust gap by transitioning from model-level scoring to continuous, system-level AgentOps evaluations, mapping real-world financial use cases directly to material risks, regulatory requirements, and measurable metrics.
Strategic Objectives
Existing Materials
Core Framework Repo: LINK
Reference Implementation: LINK (The FinSight AI Agent, a multi-agent, metacognitive system designed for earnings call analysis).
Leaderboard Repo: LINK (A framework for evaluating LLM performance across diverse financial tasks).
OSFF Toronto Training Workshop: LINK
Additional Information (Optional)
This will be a multi-repo project with ai-eval-framework holding the framework itself and the various use-case implementation contributed in their own repository as software.
Maintainer Team
Compliance & Requirements Agreement
Administrative Consent & Transfer Authorization
Socialization Acknowledgement