Skip to content

Proposed: AI Evaluation Framework - APPROVED #438

Description

@caldeirav

Contribution Prerequisites

Target Lifecycle Stage

Labs

FINOS Member Organization Name

Red Hat

Name

AI Evaluation Framework

Slug

ai-evaluation-framework

Business Problem & Solution

Business Problem

  • General-purpose AI benchmarks often fall short for the unique requirements of the finance sector.
  • Financial tasks rarely yield a single "correct" answer and require precise alignment with financial ontologies and regulatory standards.
  • Current focus on model-level evaluations do not fully capture the complexity of real-world AI systems, which combine AI and non-AI components.
  • Enterprises increasingly face a "Day 2" Reality Gap when moving AI agents into production, struggling with information overload, runaway costs, and a lack of governance and trust.
  • There is a significant risk of "safety-washing," where general capability improvements in models are misinterpreted as advancements in actual system safety.

Solution

  • Establish a common, open, and transparent evaluation suite for Generative AI and agentic applications in financial services, by utilizing a "Taxonomy-First" approach, mapping specific financial use cases (e.g., credit risk analysis) directly to material risks, and subsequently to measurable metrics.
  • The framework bridges the trust gap by transitioning from traditional MLOps to AgentOps, enabling continuous, outcome-oriented evaluation and behavioral auditing for complex AI workflows. It acts as the critical measurement layer within the FINOS Governance-as-Code pipeline, allowing organizations to continuously evaluate their models, agents, and controls against FINOS AI Governance Framework (AIGF) risks.

Mission Statement

To establish a common, open, and transparent evaluation suite for Generative AI in financial services that bridges the AI trust gap by transitioning from model-level scoring to continuous, system-level AgentOps evaluations, mapping real-world financial use cases directly to material risks, regulatory requirements, and measurable metrics.

Strategic Objectives

  • System-Level Evaluation: Move beyond generic model leaderboards to rigorously evaluate end-to-end workflows, multi-agent systems, and retrieval-augmented generation (RAG) operating in complex financial contexts.
  • Taxonomy-First Framework: Systematically align financial use cases to operational and compliance risks (such as hallucinations, bias, and non-determinism) to define clear, quantitative industry benchmarks for trustworthy AI.
  • Governance-as-Code Integration: Function as the open measurement layer in the FINOS Governance-as-Code pipeline, offering continuous auditability, metacognitive guardrails, and runtime evaluation across the system lifecycle.
  • Mutualized Innovation: Accelerate enterprise-grade AI adoption across the financial sector through open synthetic datasets, repeatable test cases, and shared reference architectures.

Existing Materials

Core Framework Repo: LINK
Reference Implementation: LINK (The FinSight AI Agent, a multi-agent, metacognitive system designed for earnings call analysis).
Leaderboard Repo: LINK (A framework for evaluating LLM performance across diverse financial tasks).
OSFF Toronto Training Workshop: LINK

Additional Information (Optional)

This will be a multi-repo project with ai-eval-framework holding the framework itself and the various use-case implementation contributed in their own repository as software.

Maintainer Team

Name Affiliation Email GitHub Username
Vincent Caldeira Red Hat vincent.caldeira@redhat.com @caldeirav
Jamie Macdonald ScottLogic jmacdonald@scottlogic.com @JamieWhitMac

Compliance & Requirements Agreement

Administrative Consent & Transfer Authorization

  • I agree to grant 'finos-admin' Admin/Owner access to the repository for security auditing and setup.
  • I authorize the transfer of this code/repository to the FINOS GitHub organization upon project acceptance.
  • I confirm I have the legal authority (individual or corporate) to grant these permissions.

Socialization Acknowledgement

  • I understand I may need to socialize this proposal to community@finos.org before a TOC review to gauge community interest.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions