Skip to content

feat: implement ethical ai tools for bias detection and fairness - #254

Merged
ooples merged 24 commits into
masterfrom
feat/us-nf-011-ethical-ai-tools
Nov 1, 2025
Merged

ooples merged 24 commits into
masterfrom
feat/us-nf-011-ethical-ai-tools

Conversation

@ooples

@ooples ooples commented Oct 30, 2025

Copy link
Copy Markdown
Owner

Summary

Implements comprehensive ethical AI tools for bias detection and fairness evaluation in AiDotNet (User Story us-nf-011).

Key Features

  • BiasDetector Class: Concrete algorithms for detecting bias in model predictions

    • Disparate impact analysis (80% rule)
    • Statistical parity difference calculation
    • Equal opportunity difference metrics
    • Per-group statistics (TPR, FPR, Precision)
    • Support for multiple protected groups
  • FairnessEvaluator Class: Comprehensive fairness metric computation

    • Demographic Parity
    • Equal Opportunity
    • Equalized Odds
    • Predictive Parity
    • Disparate Impact
    • Statistical Parity Difference
  • Integration: Seamless integration with existing interpretability infrastructure

    • Updated InterpretableModelHelper.ValidateFairnessAsync with real implementation
    • Updated VectorModel to use new fairness evaluation
    • Accessible through IInterpretableModel interface
  • Comprehensive Testing: 25+ unit tests covering:

    • Null argument handling
    • Single and multiple group scenarios
    • Balanced and unbalanced predictions
    • With and without actual labels
    • Different numeric types (double, float)
    • Edge cases (all zero predictions, etc.)

Technical Details

  • net462 Compatible: All code uses constructors instead of required keyword
  • Type Safety: Uses string keys for group dictionaries to avoid .NET 8 notnull constraint issues
  • Generic Support: Works with any numeric type T (double, float, etc.)
  • No External Dependencies: Built on existing AiDotNet infrastructure

Files Added

  • src/Interpretability/BiasDetector.cs (408 lines)
  • src/Interpretability/FairnessEvaluator.cs (472 lines)
  • tests/UnitTests/Interpretability/BiasDetectorTests.cs (232 lines)
  • tests/UnitTests/Interpretability/FairnessEvaluatorTests.cs (284 lines)

Usage Example

```csharp
// Detect bias in predictions
var detector = new BiasDetector();
var result = detector.DetectBias(predictions, sensitiveFeature, actualLabels);

if (result.HasBias)
{
Console.WriteLine(result.Message);
Console.WriteLine($"Disparate Impact: {result.DisparateImpactRatio}");
}

// Evaluate fairness metrics
var evaluator = new FairnessEvaluator();
var metrics = await evaluator.EvaluateFairnessAsync(model, inputs, sensitiveFeatureIndex, labels);

Console.WriteLine($"Demographic Parity: {metrics.DemographicParity}");
Console.WriteLine($"Equal Opportunity: {metrics.EqualOpportunity}");
```

Impact

This implementation provides the foundation for ethical AI capabilities in AiDotNet, enabling developers to:

  • Detect and measure bias in their models
  • Evaluate fairness across protected groups
  • Meet regulatory requirements (e.g., EU AI Act)
  • Build more trustworthy and accountable AI systems

🤖 Generated with Claude Code

@coderabbitai

coderabbitai Bot commented Oct 30, 2025 •

Copy link
Copy Markdown
Contributor

Note

Other AI code review bot(s) detected

CodeRabbit has detected other AI code review bot(s) in this pull request and will avoid duplicating their findings in the review comments. This may lead to a less comprehensive review.

Summary by CodeRabbit

  • New Features
    • Added bias detection capabilities for ML model predictions to identify disparate impact and demographic parity issues.
    • Added fairness evaluation tools to assess model performance across different demographic groups.
    • Model builder now supports configuration of bias detectors and fairness evaluators for ethical AI assessment.
    • Multiple bias detection methods available with configurable thresholds.

Walkthrough

Adds a fairness and bias-analysis subsystem: new generic interfaces (IBiasDetector, IFairnessEvaluator), base classes and helpers, multiple concrete detectors and evaluators, builder/result integration to carry configured components, and unit tests validating behavior across types and edge cases.

Changes

Cohort / File(s) Summary
Interfaces & Builder
src/Interfaces/IBiasDetector.cs, src/Interfaces/IFairnessEvaluator.cs, src/Interfaces/IPredictionModelBuilder.cs
Added IBiasDetector<T> and IFairnessEvaluator<T>; extended IPredictionModelBuilder with fluent ConfigureBiasDetector and ConfigureFairnessEvaluator methods.
Base classes & Result types
src/Interpretability/BiasDetectorBase.cs, src/Interpretability/FairnessEvaluatorBase.cs, src/Interpretability/BiasDetectionResult.cs, src/Interpretability/InterpretabilityMetricsHelper.cs
Added abstract bases for detectors/evaluators (validation, numeric ops, comparison helpers), BiasDetectionResult<T>, and InterpretabilityMetricsHelper<T> utilities.
Concrete bias detectors
src/Interpretability/DisparateImpactBiasDetector.cs, src/Interpretability/DemographicParityBiasDetector.cs, src/Interpretability/EqualOpportunityBiasDetector.cs
Implemented Disparate Impact, Demographic Parity, and Equal Opportunity detectors extending BiasDetectorBase<T> with thresholds and per-group computations.
Fairness evaluators
src/Interpretability/BasicFairnessEvaluator.cs, src/Interpretability/ComprehensiveFairnessEvaluator.cs, src/Interpretability/GroupFairnessEvaluator.cs
Implemented Basic, Comprehensive, and Group fairness evaluators extending FairnessEvaluatorBase<T> producing FairnessMetrics<T> and per-group AdditionalMetrics.
Builder & Result integration
src/PredictionModelBuilder.cs, src/Models/Results/PredictionModelResult.cs
PredictionModelBuilder stores configured bias/fairness components and passes them into PredictionModelResult; PredictionModelResult gained BiasDetector and FairnessEvaluator properties and constructor/deserialization support.
Tests & config
tests/UnitTests/Interpretability/BiasDetectorTests.cs, tests/UnitTests/Interpretability/FairnessEvaluatorTests.cs, .gitignore
Added unit tests covering many scenarios, numeric types, and edge cases. .gitignore updated (worktrees/, CLAUDE.md).

Sequence Diagram(s)

sequenceDiagram
    participant Builder as PredictionModelBuilder
    participant Result as PredictionModelResult
    participant Model as IFullModel
    participant Bias as IBiasDetector
    participant Fair as IFairnessEvaluator

    Builder->>Builder: ConfigureBiasDetector(detector)
    Builder->>Builder: ConfigureFairnessEvaluator(evaluator)
    Builder->>Result: Build(..., biasDetector, fairnessEvaluator)
    Note right of Result: stores configured components

    rect #EEF8F7
      Result->>Bias: DetectBias(predictions, sensitiveFeature, labels?)
      Bias->>Bias: validate inputs → GetBiasDetectionResult()
      Bias-->>Result: BiasDetectionResult<T>
    end

    rect #FEF6F0
      Result->>Fair: EvaluateFairness(model, inputs, featureIdx, labels?)
      Fair->>Fair: validate → GetFairnessMetrics()
      Fair-->>Result: FairnessMetrics<T>
    end
Loading

Estimated code review effort

🎯 4 (Complex) | ⏱️ ~45 minutes

  • Multiple new generics, numeric-operation abstractions, and control-flow validations across interpretability and builder layers.
  • Areas needing extra attention:
    • InterpretabilityMetricsHelper — correctness and division-by-zero handling for rates/precision.
    • Consistency of IsLowerBiasBetter / IsHigherFairnessBetter and IsBetterBiasScore / IsBetterFairnessScore semantics.
    • Group aggregation and AdditionalMetrics keys/population in ComprehensiveFairnessEvaluator and GroupFairnessEvaluator.
    • Serialization/deserialization and constructor changes in PredictionModelResult to ensure detectors/evaluators round-trip.
    • Threshold validation and edge-case messages in bias detectors (e.g., single-group handling, all-zero rates).

Poem

"I nibble code and count each group, 🐰
rates and ratios in my loop,
I hop through parity and score,
so models treat all groups once more.
A tiny rabbit cheers: support and proof."

Pre-merge checks and finishing touches

❌ Failed checks (1 warning)
Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 62.07% which is insufficient. The required threshold is 80.00%. You can run @coderabbitai generate docstrings to improve docstring coverage.
✅ Passed checks (2 passed)
Check name Status Explanation
Title Check ✅ Passed The PR title "feat: implement ethical ai tools for bias detection and fairness" is fully related to the primary change in the changeset and accurately summarizes the main objectives. The raw summary confirms that the PR introduces multiple bias detection classes, fairness evaluation classes, supporting infrastructure, integration points, and comprehensive tests—all of which align with implementing ethical AI tools. The title is concise, clear, and specific enough that a teammate scanning history would immediately understand this adds bias detection and fairness evaluation capabilities to the codebase.
Description Check ✅ Passed The PR description is comprehensive and clearly related to the changeset. It provides specific details about the key features implemented (BiasDetector classes with disparate impact, statistical parity, and equal opportunity metrics; FairnessEvaluator classes with multiple fairness metrics), integration points, technical approach, comprehensive testing coverage, and practical usage examples. The description is not vague or generic—it meaningfully explains what was built, why, and how it benefits users, with concrete references to the 25+ unit tests and technical considerations like net462 compatibility.
✨ Finishing touches
  • 📝 Generate docstrings
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Post copyable unit tests in a comment
  • Commit unit tests in branch feat/us-nf-011-ethical-ai-tools

📜 Recent review details

Configuration used: CodeRabbit UI

Review profile: CHILL

Plan: Pro

📥 Commits

Reviewing files that changed from the base of the PR and between 712759a and 52081e3.

📒 Files selected for processing (3)
  • src/Interpretability/DemographicParityBiasDetector.cs (1 hunks)
  • src/Interpretability/DisparateImpactBiasDetector.cs (1 hunks)
  • src/Interpretability/EqualOpportunityBiasDetector.cs (1 hunks)
🚧 Files skipped from review as they are similar to previous changes (2)
  • src/Interpretability/DemographicParityBiasDetector.cs
  • src/Interpretability/EqualOpportunityBiasDetector.cs
🧰 Additional context used
🧬 Code graph analysis (1)
src/Interpretability/DisparateImpactBiasDetector.cs (2)
src/Interpretability/DemographicParityBiasDetector.cs (1)
  • BiasDetectionResult (34-95)
src/Interpretability/InterpretabilityMetricsHelper.cs (1)
  • InterpretabilityMetricsHelper (25-288)
⏰ Context from checks skipped due to timeout of 90000ms. You can increase the timeout in your CodeRabbit configuration to a maximum of 15 minutes (900000ms). (1)
  • GitHub Check: Build All Frameworks
🔇 Additional comments (4)
src/Interpretability/DisparateImpactBiasDetector.cs (4)

22-29: Constructor validation is well-implemented.

The threshold validation correctly enforces the valid range (0 < threshold ≤ 1), and the base constructor call with isLowerBiasBetter: false is semantically appropriate for the disparate impact metric.


39-42: Input validation properly addresses the previous concern.

The length validation between predictions and sensitiveFeature has been correctly implemented, preventing potential runtime index errors that were flagged in the previous review.


44-54: Group detection and early validation are appropriate.

The requirement for at least 2 groups is correct for comparative bias detection, and the early return with a clear message prevents unnecessary computation.


56-72: Per-group statistics computation is correctly implemented.

The loop properly delegates to InterpretabilityMetricsHelper<T> for group operations, and the null-coalescing operator on line 65 safely handles potential null group values.


Comment @coderabbitai help to get the list of available commands and usage tips.

Add comprehensive bias detection and fairness evaluation tools:
- BiasDetector class with concrete algorithms for detecting bias in predictions
- FairnessEvaluator class for computing fairness metrics (demographic parity, equal opportunity, equalized odds, disparate impact)
- Integration with existing interpretability infrastructure via IInterpretableModel
- Support for multiple fairness metrics and protected groups
- Comprehensive unit tests with 25+ test cases covering various scenarios

Implements user story us-nf-011 for ethical AI capabilities.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude <noreply@anthropic.com>
@ooples
ooples force-pushed the feat/us-nf-011-ethical-ai-tools branch from 7df379f to e65b02f Compare October 30, 2025 16:52
@ooples

ooples commented Oct 30, 2025

Copy link
Copy Markdown
Owner Author

@coderabbitai review

@coderabbitai

coderabbitai Bot commented Oct 30, 2025

Copy link
Copy Markdown
Contributor
✅ Actions performed

Review triggered.

Note: CodeRabbit is an incremental review system and does not re-review already reviewed commits. This command is applicable only when automatic reviews are paused.

@ooples

ooples commented Oct 30, 2025

Copy link
Copy Markdown
Owner Author

@coderabbitai review

@coderabbitai

coderabbitai Bot commented Oct 30, 2025

Copy link
Copy Markdown
Contributor
✅ Actions performed

Review triggered.

Note: CodeRabbit is an incremental review system and does not re-review already reviewed commits. This command is applicable only when automatic reviews are paused.

ooples and others added 8 commits October 31, 2025 10:32
…d remove async

- Change FairnessEvaluator from async to synchronous (no async operations)
- Fix dictionary key type mismatch: convert T group to string before lookup
- Update FairnessEvaluatorTests to use EvaluateFairness() instead of EvaluateFairnessAsync()
- Add ConfigureBiasDetector() and ConfigureFairnessEvaluator() to IPredictionModelBuilder
- Implement Configure methods in PredictionModelBuilder for facade integration
- Update project CLAUDE.md with architectural requirements for new features

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude <noreply@anthropic.com>
Add IBiasDetector<T> interface following IFitnessCalculator pattern.
Includes DetectBias method, IsLowerBiasBetter property, and
IsBetterBiasScore comparison method for ethical AI framework.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude <noreply@anthropic.com>
…tion

Add IFairnessEvaluator<T> interface following IFitnessCalculator pattern.
Includes EvaluateFairness method, IsHigherFairnessBetter property, and
IsBetterFairnessScore comparison method for ethical AI framework.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude <noreply@anthropic.com>
Add BiasDetectorBase<T> abstract class following FitnessCalculatorBase pattern.
Implements IBiasDetector with template method pattern for bias detection.
Provides validation and comparison logic with GetBiasDetectionResult hook.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude <noreply@anthropic.com>
Add FairnessEvaluatorBase<T> abstract class following FitnessCalculatorBase pattern.
Implements IFairnessEvaluator with template method pattern for fairness evaluation.
Provides validation and comparison logic with GetFairnessMetrics hook.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude <noreply@anthropic.com>
Add InterpretabilityMetricsHelper<T> static utility class with shared methods
for bias detection and fairness evaluation. Eliminates code duplication with
GetUniqueGroups, GetGroupIndices, GetSubset, ComputePositiveRate, and metric
computation methods (TPR, FPR, Precision).

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude <noreply@anthropic.com>
…erfaces

Add using AiDotNet.Interpretability to IBiasDetector and IFairnessEvaluator
interfaces to resolve BiasDetectionResult and FairnessMetrics type references.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude <noreply@anthropic.com>

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🧹 Nitpick comments (3)
CLAUDE.md (1)

307-455: LGTM! Excellent documentation for feature integration patterns.

The new "New Feature Integration Requirements" section provides clear, actionable guidance that aligns perfectly with the patterns used in this PR (BiasDetector, FairnessEvaluator). The examples, checklist, and workflow guidance will help maintain architectural consistency.

Note: There's a minor markdown linting issue at line 390 where bold emphasis is used instead of a proper heading. Consider changing:

**NEVER duplicate utility methods across classes**

to:

#### Never Duplicate Utility Methods

**NEVER duplicate utility methods across classes**

This would satisfy the linter while maintaining the emphasis.

src/Interpretability/BiasDetector.cs (2)

78-91: Consider making bias thresholds configurable.

The thresholds are currently hardcoded:

  • 0.8 for disparate impact (documented as the 80% rule)
  • 0.1 for statistical parity difference (not explicitly documented)

While these are reasonable defaults based on fair lending guidelines, making them configurable would allow users to adjust sensitivity for different domains or regulatory requirements.

Consider adding constructor parameters:

-public BiasDetector() : base(isLowerBiasBetter: true)
+public BiasDetector(double disparateImpactThreshold = 0.8, double statisticalParityThreshold = 0.1) 
+    : base(isLowerBiasBetter: true)
 {
+    _disparateImpactThreshold = disparateImpactThreshold;
+    _statisticalParityThreshold = statisticalParityThreshold;
 }

Then use these fields in the bias detection logic:

-result.HasBias = disparateImpactValue < 0.8 || Math.Abs(statisticalParityValue) > 0.1;
+result.HasBias = disparateImpactValue < _disparateImpactThreshold || Math.Abs(statisticalParityValue) > _statisticalParityThreshold;

137-150: Consider improving message formatting.

The message concatenation at line 148 might produce awkwardly formatted output when bias was already detected:

"Bias detected: Disparate impact ratio = 0.75, Statistical parity difference = 0.12 Equal opportunity difference = 0.15."

The appended text lacks proper punctuation and could be more readable.

Consider improving the formatting:

 if (Math.Abs(eoValue) > 0.1)
 {
     result.HasBias = true;
-    result.Message += $" Equal opportunity difference = {eoValue:F3}.";
+    if (result.Message.EndsWith("."))
+    {
+        result.Message = result.Message.TrimEnd('.') + $"; Equal opportunity difference = {eoValue:F3}.";
+    }
+    else
+    {
+        result.Message += $" Equal opportunity difference = {eoValue:F3}.";
+    }
 }
📜 Review details

Configuration used: CodeRabbit UI

Review profile: CHILL

Plan: Pro

📥 Commits

Reviewing files that changed from the base of the PR and between 1f12c5b and 7b9e16e.

📒 Files selected for processing (12)
  • CLAUDE.md (1 hunks)
  • src/Interfaces/IBiasDetector.cs (1 hunks)
  • src/Interfaces/IFairnessEvaluator.cs (1 hunks)
  • src/Interfaces/IPredictionModelBuilder.cs (1 hunks)
  • src/Interpretability/BiasDetector.cs (1 hunks)
  • src/Interpretability/BiasDetectorBase.cs (1 hunks)
  • src/Interpretability/FairnessEvaluator.cs (1 hunks)
  • src/Interpretability/FairnessEvaluatorBase.cs (1 hunks)
  • src/Interpretability/InterpretabilityMetricsHelper.cs (1 hunks)
  • src/PredictionModelBuilder.cs (2 hunks)
  • tests/UnitTests/Interpretability/BiasDetectorTests.cs (1 hunks)
  • tests/UnitTests/Interpretability/FairnessEvaluatorTests.cs (1 hunks)
🧰 Additional context used
🧬 Code graph analysis (9)
src/Interfaces/IPredictionModelBuilder.cs (1)
src/Interpretability/FairnessEvaluator.cs (6)
  • T (138-144)
  • T (149-155)
  • T (160-174)
  • T (179-185)
  • T (190-202)
  • T (207-213)
src/Interpretability/BiasDetectorBase.cs (1)
src/Helpers/MathHelper.cs (2)
  • INumericOperations (33-61)
  • MathHelper (16-987)
tests/UnitTests/Interpretability/FairnessEvaluatorTests.cs (1)
src/Interpretability/FairnessEvaluator.cs (2)
  • FairnessEvaluator (13-214)
  • FairnessEvaluator (18-20)
src/Interpretability/FairnessEvaluatorBase.cs (1)
src/Helpers/MathHelper.cs (2)
  • INumericOperations (33-61)
  • MathHelper (16-987)
src/Interpretability/BiasDetector.cs (3)
src/Interpretability/FairnessEvaluator.cs (7)
  • T (138-144)
  • T (149-155)
  • T (160-174)
  • T (179-185)
  • T (190-202)
  • T (207-213)
  • Dictionary (100-133)
src/Interpretability/InterpretabilityMetricsHelper.cs (1)
  • InterpretabilityMetricsHelper (25-288)
src/Helpers/MathHelper.cs (1)
  • MathHelper (16-987)
src/PredictionModelBuilder.cs (1)
src/Interfaces/IPredictionModelBuilder.cs (12)
  • IPredictionModelBuilder (35-35)
  • IPredictionModelBuilder (51-51)
  • IPredictionModelBuilder (66-66)
  • IPredictionModelBuilder (81-81)
  • IPredictionModelBuilder (99-99)
  • IPredictionModelBuilder (133-133)
  • IPredictionModelBuilder (151-151)
  • IPredictionModelBuilder (170-170)
  • IPredictionModelBuilder (186-186)
  • IPredictionModelBuilder (305-305)
  • IPredictionModelBuilder (322-322)
  • TOutput (219-219)
tests/UnitTests/Interpretability/BiasDetectorTests.cs (1)
src/Interpretability/BiasDetector.cs (2)
  • BiasDetector (13-152)
  • BiasDetector (18-20)
src/Interpretability/FairnessEvaluator.cs (2)
src/Interpretability/InterpretabilityMetricsHelper.cs (1)
  • InterpretabilityMetricsHelper (25-288)
src/Helpers/MathHelper.cs (1)
  • MathHelper (16-987)
src/Interpretability/InterpretabilityMetricsHelper.cs (1)
src/Helpers/MathHelper.cs (2)
  • INumericOperations (33-61)
  • MathHelper (16-987)
🪛 GitHub Check: Build All Frameworks
tests/UnitTests/Interpretability/FairnessEvaluatorTests.cs

[warning] 36-36:
Cannot convert null literal to non-nullable reference type.


[warning] 24-24:
Cannot convert null literal to non-nullable reference type.

tests/UnitTests/Interpretability/BiasDetectorTests.cs

[warning] 32-32:
Cannot convert null literal to non-nullable reference type.


[warning] 21-21:
Cannot convert null literal to non-nullable reference type.

🪛 markdownlint-cli2 (0.18.1)
CLAUDE.md

390-390: Emphasis used instead of a heading

(MD036, no-emphasis-as-heading)

🔇 Additional comments (36)
src/Interfaces/IPredictionModelBuilder.cs (1)

291-322: LGTM! Excellent addition of ethical AI configuration methods.

The two new methods ConfigureBiasDetector and ConfigureFairnessEvaluator follow the established builder pattern perfectly and include comprehensive, beginner-friendly documentation. The design properly uses interface types for flexibility and maintains consistency with existing configuration methods.

src/Interpretability/InterpretabilityMetricsHelper.cs (6)

45-83: LGTM! Clean utility methods for group analysis.

The GetUniqueGroups and GetGroupIndices methods are well-implemented with proper type-safe comparisons using _numOps.Equals. The documentation clearly explains the purpose and includes helpful examples.


102-110: LGTM! Straightforward subsetting implementation.

The GetSubset method correctly creates a new vector from specified indices. The implementation is clean and efficient.


128-143: LGTM! Correct positive rate calculation.

The ComputePositiveRate method properly handles edge cases (empty predictions) and correctly counts positive predictions (> 0) with type-safe numeric operations.


164-191: LGTM! Correct TPR implementation with proper edge case handling.

The ComputeTruePositiveRate method correctly implements the TPR formula (TP / actual positives) and handles edge cases appropriately, including empty inputs and zero actual positives.


212-239: LGTM! Correct FPR implementation.

The ComputeFalsePositiveRate method correctly implements the FPR formula (FP / actual negatives) with proper edge case handling for empty inputs and zero actual negatives.


260-287: LGTM! Correct precision implementation.

The ComputePrecision method correctly implements the precision formula (TP / predicted positives) with appropriate edge case handling for empty inputs and zero predicted positives.

src/Interpretability/BiasDetectorBase.cs (3)

38-73: LGTM! Proper base class initialization.

The constructor and fields are well-designed, using MathHelper.GetNumericOperations<T>() for type-agnostic numeric operations and following .NET 4.6.2 compatibility requirements.


97-112: LGTM! Comprehensive input validation.

The DetectBias method provides thorough validation of inputs with clear error messages, properly checking for null references and length mismatches before delegating to the abstract method.


176-181: LGTM! Correct bias score comparison logic.

The IsBetterBiasScore method correctly implements the comparison logic, using the appropriate operator based on whether lower or higher bias scores indicate better fairness.

tests/UnitTests/Interpretability/FairnessEvaluatorTests.cs (5)

15-37: LGTM! Proper null argument validation tests.

The null argument tests correctly verify that ArgumentNullException is thrown for null model and null inputs. The static analysis warnings about null literals are false positives—passing null to test exception behavior is the correct approach.


39-50: LGTM! Proper bounds validation test.

The test correctly verifies that an ArgumentOutOfRangeException is thrown when an invalid sensitive feature index is provided.


52-116: LGTM! Excellent fairness detection tests.

The balanced and unbalanced prediction tests are well-designed:

  • Balanced test verifies zero bias metrics when groups are treated equally
  • Unbalanced test confirms bias detection when one group receives 100% positive predictions and the other 0%

These tests provide strong coverage of the core fairness detection functionality.


118-228: LGTM! Comprehensive edge case and feature coverage.

These tests provide excellent coverage:

  • Label-dependent metrics (Equal Opportunity, Equalized Odds, Predictive Parity)
  • Edge case of single group (correctly returns neutral metrics)
  • Additional group statistics in results
  • Multi-group scenarios (3 groups)

The test suite thoroughly validates the FairnessEvaluator functionality.


230-305: LGTM! Thorough validation of metric calculations and type support.

These tests verify:

  • Generic type support (float in addition to double)
  • Correct calculation of Demographic Parity (max - min positive rate)
  • Correct calculation of Disparate Impact (min / max positive rate)

The specific examples with expected values (0.5 for both metrics in different scenarios) provide clear regression protection.

src/Interfaces/IBiasDetector.cs (1)

1-96: LGTM! Well-designed interface with excellent documentation.

The interface design is clear, cohesive, and well-documented. The method signatures properly support the bias detection workflow, and the beginner-friendly documentation adds significant value for developers new to ethical AI concepts.

src/PredictionModelBuilder.cs (1)

334-366: Verify integration with model result and interpretability interfaces.

These fluent configuration methods correctly follow the builder pattern, but based on the PR objectives mentioning "InterpretableModelHelper.ValidateFairnessAsync implemented" and "VectorModel updated to use the new fairness evaluation," there should be code that passes these configured components to the model or exposes them through the interpretability interface.

Please confirm:

  1. How are the configured bias detector and fairness evaluator accessed by the trained model?
  2. Is there integration code in PredictionModelResult or IInterpretableModel that uses these components?
  3. Should Build() pass these to the model result for later use?
tests/UnitTests/Interpretability/BiasDetectorTests.cs (3)

14-22: LGTM! Proper null argument validation test.

This test correctly verifies that DetectBias throws ArgumentNullException when predictions are null. The static analysis warning about "Cannot convert null literal to non-nullable reference type" is a false positive in this context—passing null is intentional to test the validation logic.


24-33: LGTM! Proper null argument validation test.

This test correctly verifies that DetectBias throws ArgumentNullException when the sensitive feature is null. The static analysis warning is a false positive—passing null is intentional here to verify proper input validation.


35-229: LGTM! Comprehensive test coverage for bias detection scenarios.

The test suite thoroughly covers:

  • Edge cases (single group, all-zero predictions)
  • Balanced and unbalanced group scenarios
  • Disparate impact calculations with various thresholds
  • Additional metrics when actual labels are provided (TPR, FPR, Precision)
  • Multiple groups (non-binary sensitive features)
  • Different numeric types (double and float)
  • Group statistics validation

The test data, expected values, and assertions are all correct and well-structured.

src/Interfaces/IFairnessEvaluator.cs (1)

1-105: LGTM! Well-designed fairness evaluation interface.

The interface design is clean and well-documented. The method signature appropriately takes the model as a parameter, allowing for evaluation of fairness by generating predictions internally. The approach of using sensitiveFeatureIndex to extract the sensitive feature from the input matrix is a good design choice that keeps the interface flexible.

The comprehensive documentation explaining various fairness metrics (Demographic Parity, Equal Opportunity, Equalized Odds, Predictive Parity) is particularly helpful for developers implementing or using this interface.

src/Interpretability/FairnessEvaluator.cs (5)

13-20: LGTM! Constructor design is appropriate.

Setting isHigherFairnessBetter: false is correct for this evaluator since most of the computed metrics (Demographic Parity, Equal Opportunity, Equalized Odds, Predictive Parity, Statistical Parity Difference) are difference-based metrics where lower values (closer to 0) indicate better fairness.


25-95: LGTM! Main evaluation logic is well-structured.

The method correctly:

  • Generates predictions and extracts the sensitive feature
  • Handles the insufficient groups case gracefully by returning neutral metrics
  • Conditionally computes label-dependent metrics only when actual labels are provided
  • Attaches comprehensive per-group statistics to the result

The logic flow is clear and handles edge cases appropriately.


100-133: LGTM! Group statistics computation is correct.

The method properly iterates through groups, computes statistics using the helper methods, and conditionally computes label-dependent metrics. The use of string keys for the dictionary aligns with the PR's technical notes about .NET compatibility.


135-213: LGTM! Fairness metric calculations are correct.

All metric computations follow standard fairness definitions:

  • Demographic Parity and Statistical Parity Difference correctly compute the same value (max - min positive rate), which is accurate as these are synonymous terms in fairness literature
  • Equal Opportunity measures TPR differences
  • Equalized Odds takes the maximum of TPR and FPR differences
  • Predictive Parity measures precision differences
  • Disparate Impact properly handles division by zero

The conversions to/from double are necessary for generic type support and are handled correctly.


216-265: LGTM! Helper class is well-designed.

The GroupStatistics<T> class appropriately encapsulates per-group metrics. The constructor properly initializes all numeric fields to zero using the numeric operations abstraction, ensuring safe defaults for all supported numeric types.

The use of default(T)! on line 258 is necessary for generic type constraints and is safe in this context where T is always a numeric type.

src/Interpretability/FairnessEvaluatorBase.cs (4)

69-73: LGTM!

The constructor properly initializes the numeric operations helper and fairness direction flag. The use of MathHelper.GetNumericOperations<T>() follows the established pattern in the codebase.


98-117: LGTM!

The validation logic is comprehensive and follows best practices:

  • Null checks for required parameters
  • Range validation for the sensitive feature index
  • Length consistency validation for optional actual labels
  • Clear, informative error messages

141-145: LGTM!

The abstract method pattern provides a clean separation of concerns: the base class handles common validation while derived classes implement specific fairness metric calculations.


183-188: LGTM!

The comparison logic correctly accounts for the fairness metric direction (whether higher or lower scores indicate better fairness) and uses type-safe numeric operations.

src/Interpretability/BiasDetector.cs (6)

18-20: LGTM!

The constructor correctly initializes the base class with isLowerBiasBetter: true, which is the appropriate direction for bias metrics.


35-40: LGTM!

The early return for insufficient groups is correct—bias detection requires at least two groups for comparison. The message clearly explains why no bias analysis can be performed.


42-58: LGTM!

The group statistics computation correctly:

  • Extracts group-specific predictions
  • Computes positive rates per group
  • Tracks group sizes
  • Uses string keys for dictionaries (consistent with the PR's design to avoid .NET 8 constraints)

66-73: LGTM!

The division by zero check is properly implemented, returning a disparate impact ratio of 1.0 when the maximum rate is zero (indicating perfect parity in that edge case).


112-134: LGTM!

The additional metrics computation correctly:

  • Computes TPR, FPR, and precision for each group when actual labels are provided
  • Uses the established helper methods from InterpretabilityMetricsHelper
  • Stores results in the appropriate dictionaries

213-225: LGTM!

The constructor properly initializes all collections and numeric fields:

  • All dictionaries are instantiated to prevent null reference exceptions
  • Numeric fields are initialized to Zero using the appropriate numeric operations
  • Defensive initialization ensures the object is always in a valid state

Comment thread src/PredictionModelBuilder.cs
ooples and others added 7 commits October 31, 2025 20:52
Implements bias detection using the disparate impact ratio (80% rule).
- Ratio = min positive rate / max positive rate
- Bias detected when ratio < 0.8
- Higher ratio is better (closer to 1.0 means less bias)

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude <noreply@anthropic.com>
…ation

Implements bias detection using statistical parity difference.
- Difference = max positive rate - min positive rate
- Bias detected when |difference| > 0.1 (10%)
- Lower difference is better (0 means perfect parity)

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude <noreply@anthropic.com>
…tion

Implements bias detection using equal opportunity (TPR difference).
- Requires actual labels to compute true positive rates
- Difference = max TPR - min TPR
- Bias detected when |difference| > 0.1 (10%)
- Lower difference is better (0 means perfect equality)

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude <noreply@anthropic.com>
…tation

Computes all major fairness metrics in one evaluator:
- Demographic parity
- Equal opportunity
- Equalized odds
- Predictive parity
- Disparate impact
- Statistical parity difference
- Per-group statistics (TPR, FPR, Precision)

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude <noreply@anthropic.com>
Computes only fundamental fairness metrics:
- Demographic parity (statistical parity difference)
- Disparate impact ratio
- Does not require actual labels
- Lightweight alternative for quick fairness checks

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude <noreply@anthropic.com>
Focuses on group-level performance equity metrics:
- Equal opportunity (TPR difference)
- Equalized odds (max of TPR and FPR differences)
- Predictive parity (precision difference)
- Requires actual labels for meaningful analysis
- Ensures similar error rates across demographic groups

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude <noreply@anthropic.com>
…valuator

GroupStatistics class is now defined in ComprehensiveFairnessEvaluator.cs
Removed duplicate definition to fix CS0101 compilation error

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude <noreply@anthropic.com>
Copilot AI review requested due to automatic review settings November 1, 2025 01:33

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull Request Overview

This PR introduces ethical AI evaluation capabilities to AiDotNet by adding bias detection and fairness evaluation functionality. The implementation follows the library's established architectural patterns with interfaces, base classes, and multiple concrete implementations.

Key changes:

  • Adds bias detection infrastructure with 4 detector implementations (base, disparate impact, demographic parity, equal opportunity)
  • Adds fairness evaluation with 4 evaluator implementations (base, basic, comprehensive, group-focused)
  • Introduces shared utility helper class to avoid code duplication
  • Integrates components into PredictionModelBuilder via new configuration methods
  • Updates CLAUDE.md with architectural guidelines for future feature development

Reviewed Changes

Copilot reviewed 18 out of 18 changed files in this pull request and generated 17 comments.

Show a summary per file
File Description
FairnessEvaluatorTests.cs Comprehensive test suite for the main FairnessEvaluator implementation
BiasDetectorTests.cs Test coverage for BiasDetector with various scenarios including edge cases
PredictionModelBuilder.cs Adds configuration methods for bias detector and fairness evaluator
InterpretabilityMetricsHelper.cs Shared utility class for common fairness metric calculations
GroupFairnessEvaluator.cs Specialized evaluator focusing on group-level performance equity
FairnessEvaluatorBase.cs Abstract base class implementing common fairness evaluation logic
FairnessEvaluator.cs Comprehensive fairness evaluator computing all major metrics
EqualOpportunityBiasDetector.cs Specialized detector using equal opportunity metric
DisparateImpactBiasDetector.cs Detector implementing 80% rule for disparate impact
DemographicParityBiasDetector.cs Detector using statistical parity difference
ComprehensiveFairnessEvaluator.cs Full-featured evaluator with all fairness metrics and internal GroupStatistics helper
BiasDetectorBase.cs Abstract base class for bias detection implementations
BiasDetector.cs Main bias detector using disparate impact metrics
BasicFairnessEvaluator.cs Lightweight evaluator for fundamental metrics only
IPredictionModelBuilder.cs Interface additions for bias detector and fairness evaluator configuration
IFairnessEvaluator.cs Interface defining fairness evaluation contract
IBiasDetector.cs Interface defining bias detection contract
CLAUDE.md Documentation of architectural patterns and integration requirements

💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.

Comment thread src/Interpretability/FairnessEvaluator.cs Outdated
Comment on lines +197 to +216
internal class GroupStatistics<T>
{
public T GroupValue { get; set; }
public int Size { get; set; }
public T PositiveRate { get; set; }
public T TruePositiveRate { get; set; }
public T FalsePositiveRate { get; set; }
public T Precision { get; set; }

public GroupStatistics()
{
var numOps = MathHelper.GetNumericOperations<T>();
GroupValue = default(T)!;
Size = 0;
PositiveRate = numOps.Zero;
TruePositiveRate = numOps.Zero;
FalsePositiveRate = numOps.Zero;
Precision = numOps.Zero;
}
}

Copilot AI Nov 1, 2025

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This GroupStatistics class is duplicated in FairnessEvaluator.cs. As stated in CLAUDE.md section 3 ('NEVER duplicate utility methods across classes'), this should be extracted to a separate file and shared between both implementations.

Suggested change
internal class GroupStatistics<T>
{
public T GroupValue { get; set; }
public int Size { get; set; }
public T PositiveRate { get; set; }
public T TruePositiveRate { get; set; }
public T FalsePositiveRate { get; set; }
public T Precision { get; set; }
public GroupStatistics()
{
var numOps = MathHelper.GetNumericOperations<T>();
GroupValue = default(T)!;
Size = 0;
PositiveRate = numOps.Zero;
TruePositiveRate = numOps.Zero;
FalsePositiveRate = numOps.Zero;
Precision = numOps.Zero;
}
}
// GroupStatistics<T> is now defined in a shared file (GroupStatistics.cs).

Copilot uses AI. Check for mistakes.
Comment thread src/Interpretability/FairnessEvaluator.cs Outdated
Comment thread src/Interpretability/FairnessEvaluator.cs Outdated
Comment on lines +103 to +108
foreach (var group in groups)
{
string groupKey = group?.ToString() ?? "unknown";
fairnessMetrics.AdditionalMetrics[$"Group_{groupKey}_PositiveRate"] = groupPositiveRates[groupKey];
fairnessMetrics.AdditionalMetrics[$"Group_{groupKey}_Size"] = _numOps.FromDouble(groupSizes[groupKey]);
}

Copilot AI Nov 1, 2025

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This foreach loop immediately maps its iteration variable to another variable - consider mapping the sequence explicitly using '.Select(...)'.

Copilot uses AI. Check for mistakes.
Comment thread src/Interpretability/BiasDetector.cs Outdated
Comment on lines +68 to +75
if (_numOps.Equals(maxRate, _numOps.Zero))
{
result.DisparateImpactRatio = _numOps.One;
}
else
{
result.DisparateImpactRatio = _numOps.Divide(minRate, maxRate);
}

Copilot AI Nov 1, 2025

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Both branches of this 'if' statement write to the same variable - consider using '?' to express intent better.

Suggested change
if (_numOps.Equals(maxRate, _numOps.Zero))
{
result.DisparateImpactRatio = _numOps.One;
}
else
{
result.DisparateImpactRatio = _numOps.Divide(minRate, maxRate);
}
result.DisparateImpactRatio = _numOps.Equals(maxRate, _numOps.Zero)
? _numOps.One
: _numOps.Divide(minRate, maxRate);

Copilot uses AI. Check for mistakes.
Comment on lines +81 to +88
if (result.HasBias)
{
result.Message = $"Bias detected: Disparate impact ratio = {disparateImpactValue:F3} (below 0.8 threshold)";
}
else
{
result.Message = $"No significant bias detected: Disparate impact ratio = {disparateImpactValue:F3}";
}

Copilot AI Nov 1, 2025

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Both branches of this 'if' statement write to the same variable - consider using '?' to express intent better.

Suggested change
if (result.HasBias)
{
result.Message = $"Bias detected: Disparate impact ratio = {disparateImpactValue:F3} (below 0.8 threshold)";
}
else
{
result.Message = $"No significant bias detected: Disparate impact ratio = {disparateImpactValue:F3}";
}
result.Message = result.HasBias
? $"Bias detected: Disparate impact ratio = {disparateImpactValue:F3} (below 0.8 threshold)"
: $"No significant bias detected: Disparate impact ratio = {disparateImpactValue:F3}";

Copilot uses AI. Check for mistakes.
Comment on lines +73 to +80
if (result.HasBias)
{
result.Message = $"Bias detected: Statistical parity difference = {statisticalParityValue:F3} (exceeds 0.1 threshold)";
}
else
{
result.Message = $"No significant bias detected: Statistical parity difference = {statisticalParityValue:F3}";
}

Copilot AI Nov 1, 2025

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Both branches of this 'if' statement write to the same variable - consider using '?' to express intent better.

Suggested change
if (result.HasBias)
{
result.Message = $"Bias detected: Statistical parity difference = {statisticalParityValue:F3} (exceeds 0.1 threshold)";
}
else
{
result.Message = $"No significant bias detected: Statistical parity difference = {statisticalParityValue:F3}";
}
result.Message = result.HasBias
? $"Bias detected: Statistical parity difference = {statisticalParityValue:F3} (exceeds 0.1 threshold)"
: $"No significant bias detected: Statistical parity difference = {statisticalParityValue:F3}";

Copilot uses AI. Check for mistakes.
Comment on lines +91 to +98
if (result.HasBias)
{
result.Message = $"Bias detected: Equal opportunity difference = {eoValue:F3} (exceeds 0.1 threshold)";
}
else
{
result.Message = $"No significant bias detected: Equal opportunity difference = {eoValue:F3}";
}

Copilot AI Nov 1, 2025

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Both branches of this 'if' statement write to the same variable - consider using '?' to express intent better.

Suggested change
if (result.HasBias)
{
result.Message = $"Bias detected: Equal opportunity difference = {eoValue:F3} (exceeds 0.1 threshold)";
}
else
{
result.Message = $"No significant bias detected: Equal opportunity difference = {eoValue:F3}";
}
result.Message = result.HasBias
? $"Bias detected: Equal opportunity difference = {eoValue:F3} (exceeds 0.1 threshold)"
: $"No significant bias detected: Equal opportunity difference = {eoValue:F3}";

Copilot uses AI. Check for mistakes.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 5

🧹 Nitpick comments (4)
src/Interpretability/DisparateImpactBiasDetector.cs (1)

67-75: Consider adding a warning when all predictions are zero.

When maxRate is zero, the code sets the ratio to 1.0, indicating no bias. While technically correct (all groups are treated equally), this might mask a broken model where no positive predictions are made at all.

Consider this enhancement to provide more context:

-            // Avoid division by zero
             if (_numOps.Equals(maxRate, _numOps.Zero))
             {
                 result.DisparateImpactRatio = _numOps.One;
+                result.Message = "All groups have zero positive predictions. Disparate impact ratio set to 1.0.";
+                result.HasBias = false;
+                return result;
             }
             else
             {
                 result.DisparateImpactRatio = _numOps.Divide(minRate, maxRate);
             }
src/Interpretability/EqualOpportunityBiasDetector.cs (2)

44-66: Consider clarifying the comment at line 44.

The comment states "Compute positive prediction rates for each group" but the loop also computes TPRs (when actual labels are provided) and group sizes. While the positive rates are stored in the result for inspection, they're not used in the Equal Opportunity calculation itself—only the TPRs matter for this metric.

Consider updating the comment to be more comprehensive:

-// Compute positive prediction rates for each group
+// Compute per-group statistics: positive rates, sizes, and TPRs (if labels provided)

80-98: Minor optimization: Math.Abs() is redundant.

Since the Equal Opportunity difference is computed as max TPR - min TPR (line 85), the result is always non-negative. The Math.Abs() at line 89 is therefore unnecessary, though it doesn't affect correctness.

Consider simplifying the threshold check:

-result.HasBias = Math.Abs(eoValue) > 0.1;
+result.HasBias = eoValue > 0.1;

Alternatively, if you want to be defensive against future refactoring, the current code is fine.

src/Interpretability/DemographicParityBiasDetector.cs (1)

69-83: LGTM! Clear bias detection logic.

The bias threshold check and messaging are well-implemented. The 0.1 (10%) threshold aligns with standard fairness analysis practices.

Optional refinement: Since statisticalParityDifference is computed as maxRate - minRate (line 67), it's always non-negative. The Math.Abs() on line 71 is redundant but harmless—removing it would be a minor optimization.

Optional: Remove redundant Math.Abs()
-            result.HasBias = Math.Abs(statisticalParityValue) > 0.1;
+            result.HasBias = statisticalParityValue > 0.1;
📜 Review details

Configuration used: CodeRabbit UI

Review profile: CHILL

Plan: Pro

📥 Commits

Reviewing files that changed from the base of the PR and between 7b9e16e and 9be32f2.

📒 Files selected for processing (7)
  • src/Interpretability/BasicFairnessEvaluator.cs (1 hunks)
  • src/Interpretability/ComprehensiveFairnessEvaluator.cs (1 hunks)
  • src/Interpretability/DemographicParityBiasDetector.cs (1 hunks)
  • src/Interpretability/DisparateImpactBiasDetector.cs (1 hunks)
  • src/Interpretability/EqualOpportunityBiasDetector.cs (1 hunks)
  • src/Interpretability/FairnessEvaluator.cs (1 hunks)
  • src/Interpretability/GroupFairnessEvaluator.cs (1 hunks)
🧰 Additional context used
🧬 Code graph analysis (7)
src/Interpretability/DisparateImpactBiasDetector.cs (1)
src/Interpretability/InterpretabilityMetricsHelper.cs (1)
  • InterpretabilityMetricsHelper (25-288)
src/Interpretability/DemographicParityBiasDetector.cs (1)
src/Interpretability/InterpretabilityMetricsHelper.cs (1)
  • InterpretabilityMetricsHelper (25-288)
src/Interpretability/BasicFairnessEvaluator.cs (1)
src/Interpretability/InterpretabilityMetricsHelper.cs (1)
  • InterpretabilityMetricsHelper (25-288)
src/Interpretability/ComprehensiveFairnessEvaluator.cs (2)
src/Interpretability/InterpretabilityMetricsHelper.cs (1)
  • InterpretabilityMetricsHelper (25-288)
src/Helpers/MathHelper.cs (1)
  • MathHelper (16-987)
src/Interpretability/GroupFairnessEvaluator.cs (1)
src/Interpretability/InterpretabilityMetricsHelper.cs (1)
  • InterpretabilityMetricsHelper (25-288)
src/Interpretability/FairnessEvaluator.cs (1)
src/Interpretability/InterpretabilityMetricsHelper.cs (1)
  • InterpretabilityMetricsHelper (25-288)
src/Interpretability/EqualOpportunityBiasDetector.cs (1)
src/Interpretability/InterpretabilityMetricsHelper.cs (1)
  • InterpretabilityMetricsHelper (25-288)
⏰ Context from checks skipped due to timeout of 90000ms. You can increase the timeout in your CodeRabbit configuration to a maximum of 15 minutes (900000ms). (2)
  • GitHub Check: CodeQL analysis (csharp)
  • GitHub Check: Build All Frameworks
🔇 Additional comments (9)
src/Interpretability/DisparateImpactBiasDetector.cs (2)

14-22: LGTM!

The constructor correctly initializes the base class with isLowerBiasBetter: false, which is appropriate since a higher Disparate Impact Ratio (closer to 1.0) indicates less bias.


44-60: LGTM!

The per-group rate computation is correct. The string key conversion at line 53 is consistent with the PR's design decision to use string keys for .NET 4.6.2 compatibility.

src/Interpretability/EqualOpportunityBiasDetector.cs (4)

8-22: LGTM! Well-documented class with correct initialization.

The class documentation clearly describes the Equal Opportunity metric, threshold, and requirement for actual labels. The constructor correctly initializes the base with isLowerBiasBetter: true since a difference of 0 represents perfect equality.


27-42: LGTM! Proper validation and early return pattern.

The method correctly validates that at least 2 groups exist before proceeding with bias detection, returning early with a clear message if insufficient groups are found.


68-78: LGTM! Proper handling of missing actual labels.

The code correctly validates that actual labels are provided before attempting to compute Equal Opportunity, returning early with a clear, actionable message if they're missing. This is essential since TPR calculation requires ground truth labels.


1-103: Overall: Solid implementation of Equal Opportunity bias detection.

The detector correctly implements the Equal Opportunity metric by computing TPR differences across groups. Key strengths:

  • ✅ Proper validation of inputs (group count, actual labels requirement)
  • ✅ Correct use of helper methods for group extraction and metric computation
  • ✅ Clear, actionable error messages
  • ✅ Consistent with the broader bias detection framework
  • ✅ Well-documented with XML comments

The implementation is production-ready with only minor, optional improvements suggested above.

src/Interpretability/DemographicParityBiasDetector.cs (3)

19-22: LGTM! Correct initialization for demographic parity.

The constructor correctly sets isLowerBiasBetter: true since a statistical parity difference of 0 represents perfect fairness, and any deviation from 0 indicates increasing bias.


32-42: LGTM! Proper validation for bias detection.

The code correctly validates that at least 2 groups are present before proceeding with bias analysis. The early return with a clear message is good error handling.


44-67: LGTM! Correct implementation of statistical parity calculation.

The code properly:

  • Computes positive prediction rates for each group using helper methods
  • Stores group-level metrics for transparency
  • Calculates statistical parity difference as the gap between max and min rates
  • Uses string keys for dictionaries (a documented design decision for framework compatibility)

The group identification uses actual T values for comparison, with string conversion only for dictionary keys, ensuring correctness.

Comment on lines +62 to +108
foreach (var group in groups)
{
var groupIndices = InterpretabilityMetricsHelper<T>.GetGroupIndices(sensitiveFeature, group);
var groupPredictions = InterpretabilityMetricsHelper<T>.GetSubset(predictions, groupIndices);
var positiveRate = InterpretabilityMetricsHelper<T>.ComputePositiveRate(groupPredictions);
string groupKey = group?.ToString() ?? "unknown";

groupPositiveRates[groupKey] = positiveRate;
groupSizes[groupKey] = groupIndices.Count;
}

// Compute demographic parity (statistical parity difference)
var orderedRates = groupPositiveRates.Values.OrderBy(r => Convert.ToDouble(r)).ToList();
T minRate = orderedRates.First();
T maxRate = orderedRates.Last();

T demographicParity = _numOps.Subtract(maxRate, minRate);

// Compute disparate impact
T disparateImpact;
if (_numOps.Equals(maxRate, _numOps.Zero))
{
disparateImpact = _numOps.One;
}
else
{
disparateImpact = _numOps.Divide(minRate, maxRate);
}

var fairnessMetrics = new FairnessMetrics<T>(
demographicParity: demographicParity,
equalOpportunity: _numOps.Zero, // Not computed in basic evaluator
equalizedOdds: _numOps.Zero, // Not computed in basic evaluator
predictiveParity: _numOps.Zero, // Not computed in basic evaluator
disparateImpact: disparateImpact,
statisticalParityDifference: demographicParity)
{
SensitiveFeatureIndex = sensitiveFeatureIndex
};

// Add per-group basic metrics
foreach (var group in groups)
{
string groupKey = group?.ToString() ?? "unknown";
fairnessMetrics.AdditionalMetrics[$"Group_{groupKey}_PositiveRate"] = groupPositiveRates[groupKey];
fairnessMetrics.AdditionalMetrics[$"Group_{groupKey}_Size"] = _numOps.FromDouble(groupSizes[groupKey]);
}

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🔴 Critical

Avoid collapsing distinct groups under the "unknown" label.

Here the positive-rate and size maps rely on group?.ToString() ?? "unknown" as the key. As soon as the dataset contains a real "unknown" group (or multiple null-like labels), you overwrite the earlier entry, which means the demographic parity and disparate-impact calculations operate on fewer groups than actually exist. The resulting fairness metrics are wrong.

Please switch to collision-proof keys (index + label) when you populate groupPositiveRates / groupSizes and when you emit AdditionalMetrics. For example:

-            foreach (var group in groups)
-            {
-                string groupKey = group?.ToString() ?? "unknown";
-                groupPositiveRates[groupKey] = positiveRate;
-                groupSizes[groupKey] = groupIndices.Count;
+            for (int i = 0; i < groups.Count; i++)
+            {
+                var group = groups[i];
+                string groupKey = BuildGroupKey(group, i);
+                groupPositiveRates[groupKey] = positiveRate;
+                groupSizes[groupKey] = groupIndices.Count;
             }

and

-            foreach (var group in groups)
-            {
-                string groupKey = group?.ToString() ?? "unknown";
-                fairnessMetrics.AdditionalMetrics[$"Group_{groupKey}_PositiveRate"] = groupPositiveRates[groupKey];
+            for (int i = 0; i < groups.Count; i++)
+            {
+                var group = groups[i];
+                string groupKey = BuildGroupKey(group, i);
+                string metricPrefix = $"Group_{i}_{group?.ToString() ?? "null"}";
+                fairnessMetrics.AdditionalMetrics[$"{metricPrefix}_PositiveRate"] = groupPositiveRates[groupKey];

Introduce the same BuildGroupKey helper used elsewhere so every evaluator shares a consistent, safe approach.

Committable suggestion skipped: line range outside the PR's diff.

🤖 Prompt for AI Agents
In src/Interpretability/BasicFairnessEvaluator.cs around lines 62 to 108, the
code uses group?.ToString() ?? "unknown" as the dictionary key which allows
collisions when a real "unknown" label or multiple nulls exist; change to a
collision-proof key that combines the group index and label (e.g.,
$"{index}:{label}") by using the shared BuildGroupKey helper used elsewhere so
keys are unique; use BuildGroupKey(sensitiveFeatureIndex, index, group) (or the
helper's required signature) when populating groupPositiveRates and groupSizes
and the same BuildGroupKey when writing AdditionalMetrics so keys remain
consistent and distinct.

Comment on lines +80 to +132
// Add per-group metrics to additional metrics
foreach (var group in groups)
{
string groupKey = group?.ToString() ?? "unknown";
var stats = groupStats[groupKey];
fairnessMetrics.AdditionalMetrics[$"Group_{groupKey}_PositiveRate"] = stats.PositiveRate;
fairnessMetrics.AdditionalMetrics[$"Group_{groupKey}_Size"] = _numOps.FromDouble(stats.Size);

if (actualLabels != null)
{
fairnessMetrics.AdditionalMetrics[$"Group_{groupKey}_TPR"] = stats.TruePositiveRate;
fairnessMetrics.AdditionalMetrics[$"Group_{groupKey}_FPR"] = stats.FalsePositiveRate;
fairnessMetrics.AdditionalMetrics[$"Group_{groupKey}_Precision"] = stats.Precision;
}
}

return fairnessMetrics;
}

private Dictionary<string, GroupStatistics<T>> ComputeGroupStatistics(
Vector<T> predictions,
Vector<T> sensitiveFeature,
Vector<T>? actualLabels,
List<T> groups)
{
var groupStats = new Dictionary<string, GroupStatistics<T>>();

foreach (var group in groups)
{
var groupIndices = InterpretabilityMetricsHelper<T>.GetGroupIndices(sensitiveFeature, group);
var groupPredictions = InterpretabilityMetricsHelper<T>.GetSubset(predictions, groupIndices);
string groupKey = group?.ToString() ?? "unknown";

var stats = new GroupStatistics<T>
{
GroupValue = group,
Size = groupIndices.Count,
PositiveRate = InterpretabilityMetricsHelper<T>.ComputePositiveRate(groupPredictions)
};

if (actualLabels != null)
{
var groupActualLabels = InterpretabilityMetricsHelper<T>.GetSubset(actualLabels, groupIndices);
stats.TruePositiveRate = InterpretabilityMetricsHelper<T>.ComputeTruePositiveRate(groupPredictions, groupActualLabels);
stats.FalsePositiveRate = InterpretabilityMetricsHelper<T>.ComputeFalsePositiveRate(groupPredictions, groupActualLabels);
stats.Precision = InterpretabilityMetricsHelper<T>.ComputePrecision(groupPredictions, groupActualLabels);
}

groupStats[groupKey] = stats;
}

return groupStats;
}

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🔴 Critical

Comprehensive evaluator has the same group-key overwrite bug.

This class duplicates the FairnessEvaluator logic verbatim, so the same "unknown" key collision occurs: null-valued sensitive attributes and a genuine "unknown" bucket (or any two values with identical ToString() output) overwrite each other in groupStats, leading to missing cohorts and incorrect fairness numbers.

Please apply the same fix here — generate unique keys that include the group index (and reuse that key everywhere you interact with groupStats / AdditionalMetrics). A helper like BuildGroupKey keeps the code tidy:

-            foreach (var group in groups)
-            {
-                string groupKey = group?.ToString() ?? "unknown";
-                var stats = groupStats[groupKey];
-                fairnessMetrics.AdditionalMetrics[$"Group_{groupKey}_PositiveRate"] = stats.PositiveRate;
+            for (int i = 0; i < groups.Count; i++)
+            {
+                var group = groups[i];
+                string groupKey = BuildGroupKey(group, i);
+                var stats = groupStats[groupKey];
+                string metricPrefix = $"Group_{i}_{group?.ToString() ?? "null"}";
+                fairnessMetrics.AdditionalMetrics[$"{metricPrefix}_PositiveRate"] = stats.PositiveRate;

and likewise in ComputeGroupStatistics. Ensuring every evaluator shares the collision-free key strategy restores correctness across the board.

Committable suggestion skipped: line range outside the PR's diff.

Comment thread src/Interpretability/DisparateImpactBiasDetector.cs
Comment thread src/Interpretability/FairnessEvaluator.cs Outdated
Comment on lines +64 to +142
foreach (var group in groups)
{
var groupIndices = InterpretabilityMetricsHelper<T>.GetGroupIndices(sensitiveFeature, group);
var groupPredictions = InterpretabilityMetricsHelper<T>.GetSubset(predictions, groupIndices);
string groupKey = group?.ToString() ?? "unknown";

groupSizes[groupKey] = groupIndices.Count;

if (actualLabels != null)
{
var groupActualLabels = InterpretabilityMetricsHelper<T>.GetSubset(actualLabels, groupIndices);

// Compute TPR (True Positive Rate)
var tpr = InterpretabilityMetricsHelper<T>.ComputeTruePositiveRate(groupPredictions, groupActualLabels);
groupTPRs[groupKey] = tpr;

// Compute FPR (False Positive Rate)
var fpr = InterpretabilityMetricsHelper<T>.ComputeFalsePositiveRate(groupPredictions, groupActualLabels);
groupFPRs[groupKey] = fpr;

// Compute Precision
var precision = InterpretabilityMetricsHelper<T>.ComputePrecision(groupPredictions, groupActualLabels);
groupPrecisions[groupKey] = precision;
}
}

// Compute group fairness metrics
T equalOpportunity = _numOps.Zero;
T equalizedOdds = _numOps.Zero;
T predictiveParity = _numOps.Zero;

if (actualLabels != null && groupTPRs.Count >= 2)
{
// Equal Opportunity: max TPR difference
var orderedTPRs = groupTPRs.Values.OrderBy(r => Convert.ToDouble(r)).ToList();
T minTPR = orderedTPRs.First();
T maxTPR = orderedTPRs.Last();
equalOpportunity = _numOps.Subtract(maxTPR, minTPR);

// Equalized Odds: max of TPR and FPR differences
var orderedFPRs = groupFPRs.Values.OrderBy(r => Convert.ToDouble(r)).ToList();
T minFPR = orderedFPRs.First();
T maxFPR = orderedFPRs.Last();

double tprDiff = Convert.ToDouble(_numOps.Subtract(maxTPR, minTPR));
double fprDiff = Convert.ToDouble(_numOps.Subtract(maxFPR, minFPR));
equalizedOdds = _numOps.FromDouble(Math.Max(tprDiff, fprDiff));

// Predictive Parity: max precision difference
var orderedPrecisions = groupPrecisions.Values.OrderBy(r => Convert.ToDouble(r)).ToList();
T minPrecision = orderedPrecisions.First();
T maxPrecision = orderedPrecisions.Last();
predictiveParity = _numOps.Subtract(maxPrecision, minPrecision);
}

var fairnessMetrics = new FairnessMetrics<T>(
demographicParity: _numOps.Zero, // Not primary focus in group evaluator
equalOpportunity: equalOpportunity,
equalizedOdds: equalizedOdds,
predictiveParity: predictiveParity,
disparateImpact: _numOps.One, // Not primary focus in group evaluator
statisticalParityDifference: _numOps.Zero)
{
SensitiveFeatureIndex = sensitiveFeatureIndex
};

// Add per-group performance metrics
foreach (var group in groups)
{
string groupKey = group?.ToString() ?? "unknown";
fairnessMetrics.AdditionalMetrics[$"Group_{groupKey}_Size"] = _numOps.FromDouble(groupSizes[groupKey]);

if (actualLabels != null)
{
fairnessMetrics.AdditionalMetrics[$"Group_{groupKey}_TPR"] = groupTPRs[groupKey];
fairnessMetrics.AdditionalMetrics[$"Group_{groupKey}_FPR"] = groupFPRs[groupKey];
fairnessMetrics.AdditionalMetrics[$"Group_{groupKey}_Precision"] = groupPrecisions[groupKey];
}
}

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🔴 Critical

Prevent group metric overwrites caused by "unknown" keys.

The GroupFairnessEvaluator keeps per-group TPR/FPR/Precision in dictionaries keyed by group?.ToString() ?? "unknown", and the same key is reused when filling AdditionalMetrics. If a dataset contains both a real "unknown" group and actual nulls (or just two distinct values whose ToString() collide), the later insert overwrites the prior entry and drops one cohort from the fairness calculation. That corrupts equal opportunity / equalized odds output.

Please make the group key unambiguous — e.g. incorporate the loop index into the key when you populate groupTPRs, groupFPRs, groupPrecisions, groupSizes, and when you emit the AdditionalMetrics:

-            foreach (var group in groups)
-            {
-                string groupKey = group?.ToString() ?? "unknown";
-                groupSizes[groupKey] = groupIndices.Count;
+            for (int i = 0; i < groups.Count; i++)
+            {
+                var group = groups[i];
+                string groupKey = BuildGroupKey(group, i);
+                groupSizes[groupKey] = groupIndices.Count;
                 ...
-                    var groupActualLabels = ...
-                    groupTPRs[groupKey] = tpr;
+                    var groupActualLabels = ...
+                    groupTPRs[groupKey] = tpr;
                     ...
             }

and when emitting metrics:

-            foreach (var group in groups)
-            {
-                string groupKey = group?.ToString() ?? "unknown";
-                fairnessMetrics.AdditionalMetrics[$"Group_{groupKey}_Size"] = ...
+            for (int i = 0; i < groups.Count; i++)
+            {
+                var group = groups[i];
+                string groupKey = BuildGroupKey(group, i);
+                string metricPrefix = $"Group_{i}_{group?.ToString() ?? "null"}";
+                fairnessMetrics.AdditionalMetrics[$"{metricPrefix}_Size"] = ...

Add a private BuildGroupKey helper (same as in FairnessEvaluator) so every evaluator shares the same collision-free logic.

Committable suggestion skipped: line range outside the PR's diff.

🤖 Prompt for AI Agents
In src/Interpretability/GroupFairnessEvaluator.cs around lines 64-142, the
per-group dictionaries and emitted AdditionalMetrics use group?.ToString() ??
"unknown" which can collide and overwrite metrics; add a private
BuildGroupKey(object group, int index) helper (matching FairnessEvaluator) that
returns an unambiguous key (e.g., $"{index}:{group?.ToString() ?? "unknown"}"),
use it when populating groupSizes, groupTPRs, groupFPRs, groupPrecisions
(replace string groupKey = ... with BuildGroupKey(group, i) where i is the loop
index), and use the same BuildGroupKey when emitting AdditionalMetrics so keys
are collision-free and consistent across population and output.

ooples and others added 3 commits October 31, 2025 22:07
… files

- Revert CLAUDE.md changes (project config should not be in PRs)
- Add CLAUDE.md to .gitignore to prevent future commits
- Delete BiasDetector.cs (duplicate of DisparateImpactBiasDetector)
- Delete FairnessEvaluator.cs (duplicate of ComprehensiveFairnessEvaluator)

Use the specific concrete implementations instead:
- DisparateImpactBiasDetector, DemographicParityBiasDetector, EqualOpportunityBiasDetector
- ComprehensiveFairnessEvaluator, BasicFairnessEvaluator, GroupFairnessEvaluator

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude <noreply@anthropic.com>
- Create BiasDetectionResult.cs as standalone class
- Update BiasDetectorTests to use DisparateImpactBiasDetector
- Update FairnessEvaluatorTests to use ComprehensiveFairnessEvaluator
- All tests now use concrete implementations instead of deleted base classes

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude <noreply@anthropic.com>
…edictionmodelresult

- Add BiasDetector and FairnessEvaluator properties to PredictionModelResult
- Add optional parameters to PredictionModelResult constructor
- Pass configured _biasDetector and _fairnessEvaluator from Build() method
- Users can now access bias detection and fairness evaluation tools from trained models
- Maintains nullability - only populated if configured via builder

Usage example:

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude <noreply@anthropic.com>

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 0

Caution

Some comments are outside the diff and can’t be posted inline due to platform limitations.

⚠️ Outside diff range comments (1)
src/Models/Results/PredictionModelResult.cs (1)

419-427: Restore fairness components during deserialization.

Deserialize copies Model, OptimizationResult, NormalizationInfo, and ModelMetaData from the deserialized instance but never assigns BiasDetector or FairnessEvaluator. Any persisted model that originally carried ethical-AI components will lose them after deserialization (e.g., when reloading from disk via LoadFromFile or LoadModel), so fairness/bias analysis silently stops working. Please copy these properties as well.

             if (deserializedObject != null)
             {
                 Model = deserializedObject.Model;
                 OptimizationResult = deserializedObject.OptimizationResult;
                 NormalizationInfo = deserializedObject.NormalizationInfo;
                 ModelMetaData = deserializedObject.ModelMetaData;
+                BiasDetector = deserializedObject.BiasDetector;
+                FairnessEvaluator = deserializedObject.FairnessEvaluator;
             }
🧹 Nitpick comments (1)
.gitignore (1)

370-370: Verify the directory name change from .worktrees/ to worktrees/.

The pattern changed from .worktrees/ (hidden directory) to worktrees/ (non-hidden directory). These match different directory names. Confirm this intentionally reflects the actual directory structure created by Claude Code.

📜 Review details

Configuration used: CodeRabbit UI

Review profile: CHILL

Plan: Pro

📥 Commits

Reviewing files that changed from the base of the PR and between 9be32f2 and 2c80f03.

📒 Files selected for processing (6)
  • .gitignore (1 hunks)
  • src/Interpretability/BiasDetectionResult.cs (1 hunks)
  • src/Models/Results/PredictionModelResult.cs (3 hunks)
  • src/PredictionModelBuilder.cs (3 hunks)
  • tests/UnitTests/Interpretability/BiasDetectorTests.cs (1 hunks)
  • tests/UnitTests/Interpretability/FairnessEvaluatorTests.cs (1 hunks)
🚧 Files skipped from review as they are similar to previous changes (1)
  • tests/UnitTests/Interpretability/FairnessEvaluatorTests.cs
🧰 Additional context used
🧬 Code graph analysis (4)
tests/UnitTests/Interpretability/BiasDetectorTests.cs (1)
tests/UnitTests/Interpretability/FairnessEvaluatorTests.cs (12)
  • Fact (15-25)
  • Fact (27-37)
  • Fact (39-50)
  • Fact (52-84)
  • Fact (86-116)
  • Fact (118-146)
  • Fact (148-171)
  • Fact (173-197)
  • Fact (199-228)
  • Fact (230-250)
  • Fact (252-278)
  • Fact (280-305)
src/PredictionModelBuilder.cs (2)
src/Models/Results/PredictionModelResult.cs (4)
  • PredictionModelResult (208-219)
  • PredictionModelResult (246-248)
  • PredictionModelResult (555-578)
  • TOutput (313-329)
src/Interfaces/IPredictionModelBuilder.cs (12)
  • TOutput (219-219)
  • IPredictionModelBuilder (35-35)
  • IPredictionModelBuilder (51-51)
  • IPredictionModelBuilder (66-66)
  • IPredictionModelBuilder (81-81)
  • IPredictionModelBuilder (99-99)
  • IPredictionModelBuilder (133-133)
  • IPredictionModelBuilder (151-151)
  • IPredictionModelBuilder (170-170)
  • IPredictionModelBuilder (186-186)
  • IPredictionModelBuilder (305-305)
  • IPredictionModelBuilder (322-322)
src/Interpretability/BiasDetectionResult.cs (1)
src/Helpers/MathHelper.cs (1)
  • MathHelper (16-987)
src/Models/Results/PredictionModelResult.cs (2)
src/Models/NormalizationInfo.cs (3)
  • NormalizationInfo (31-150)
  • NormalizationInfo (130-138)
  • NormalizationInfo (146-149)
src/Models/Results/OptimizationResult.cs (4)
  • OptimizationResult (37-761)
  • OptimizationResult (445-458)
  • OptimizationResult (464-481)
  • OptimizationResult (488-505)
⏰ Context from checks skipped due to timeout of 90000ms. You can increase the timeout in your CodeRabbit configuration to a maximum of 15 minutes (900000ms). (1)
  • GitHub Check: Build All Frameworks
🔇 Additional comments (1)
.gitignore (1)

371-372: Reasonable addition for Claude Code configuration hygiene.

Ignoring CLAUDE.md aligns with keeping tool-specific artifacts out of version control, consistent with the .claude/ entry at line 366.

@coderabbitai

coderabbitai Bot commented Nov 1, 2025

Copy link
Copy Markdown
Contributor
✅ Actions performed

Initiated docstring generation; will generate only if new commits exist.

@coderabbitai

coderabbitai Bot commented Nov 1, 2025 •

Copy link
Copy Markdown
Contributor

Caution

Docstrings generation - FAILED

An unexpected error occurred while opening a pull request: Repository rule violations found

Changes must be made through a pull request.

ooples and others added 4 commits October 31, 2025 23:24
…erialization

- Add BiasDetector and FairnessEvaluator to deserialization logic
- Previously these properties were lost when loading model from disk
- Now persisted models maintain their ethical AI components

Addresses code review comment about missing deserialization.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude <noreply@anthropic.com>
…ng to disparateimpactbiasdetector

- Add configurable threshold parameter (default 0.8 for 80% rule)
- Add validation for threshold range (0 < threshold <= 1)
- Use threshold in bias detection instead of hardcoded 0.8
- Add warning message when all predictions are zero
- Show threshold value in bias detection message

Addresses code review comments about configurability and edge cases.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude <noreply@anthropic.com>
…etector

- Add configurable threshold parameter (default 0.1 for 10% difference)
- Add validation for threshold range (0 < threshold <= 1)
- Use threshold in bias detection instead of hardcoded 0.1
- Show threshold value in bias detection message

Addresses code review comments about configurability.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude <noreply@anthropic.com>
….abs in equalopportunitybiasdetector

- Add configurable threshold parameter (default 0.1 for 10% difference)
- Add validation for threshold range (0 < threshold <= 1)
- Use threshold in bias detection instead of hardcoded 0.1
- Remove redundant Math.Abs() since max - min is always non-negative
- Show threshold value in bias detection message

Addresses code review comments about configurability and optimization.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude <noreply@anthropic.com)
Copilot AI review requested due to automatic review settings November 1, 2025 04:00

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

♻️ Duplicate comments (1)
src/Interpretability/DisparateImpactBiasDetector.cs (1)

34-38: Add input validation for vector lengths.

As previously noted, the method lacks validation that predictions and sensitiveFeature have matching lengths, which can cause runtime errors when accessing group indices.

🧹 Nitpick comments (3)
src/Interpretability/DemographicParityBiasDetector.cs (1)

80-87: Consider using ternary operator for message assignment.

Both branches assign to result.Message. A ternary operator would be more concise.

Apply this diff:

-    if (result.HasBias)
-    {
-        result.Message = $"Bias detected: Statistical parity difference = {statisticalParityValue:F3} (exceeds {_threshold} threshold)";
-    }
-    else
-    {
-        result.Message = $"No significant bias detected: Statistical parity difference = {statisticalParityValue:F3}";
-    }
+    result.Message = result.HasBias
+        ? $"Bias detected: Statistical parity difference = {statisticalParityValue:F3} (exceeds {_threshold} threshold)"
+        : $"No significant bias detected: Statistical parity difference = {statisticalParityValue:F3}";
src/Interpretability/DisparateImpactBiasDetector.cs (2)

75-85: Consider simplifying division-by-zero handling with ternary operator.

The if-else block can be condensed to a single ternary expression.

Apply this diff:

-    // Avoid division by zero
-    if (_numOps.Equals(maxRate, _numOps.Zero))
-    {
-        result.DisparateImpactRatio = _numOps.One;
-        result.Message = "All groups have zero positive predictions. Disparate impact ratio set to 1.0.";
-        result.HasBias = false;
-        return result;
-    }
-    else
-    {
-        result.DisparateImpactRatio = _numOps.Divide(minRate, maxRate);
-    }
+    // Avoid division by zero
+    if (_numOps.Equals(maxRate, _numOps.Zero))
+    {
+        result.DisparateImpactRatio = _numOps.One;
+        result.Message = "All groups have zero positive predictions. Disparate impact ratio set to 1.0.";
+        result.HasBias = false;
+        return result;
+    }
+
+    result.DisparateImpactRatio = _numOps.Divide(minRate, maxRate);

Note: While the early return makes a ternary less practical here, removing the else improves readability.


91-98: Consider using ternary operator for message assignment.

Both branches assign to result.Message. A ternary operator would be more concise.

Apply this diff:

-    if (result.HasBias)
-    {
-        result.Message = $"Bias detected: Disparate impact ratio = {disparateImpactValue:F3} (below {_threshold} threshold)";
-    }
-    else
-    {
-        result.Message = $"No significant bias detected: Disparate impact ratio = {disparateImpactValue:F3}";
-    }
+    result.Message = result.HasBias
+        ? $"Bias detected: Disparate impact ratio = {disparateImpactValue:F3} (below {_threshold} threshold)"
+        : $"No significant bias detected: Disparate impact ratio = {disparateImpactValue:F3}";
📜 Review details

Configuration used: CodeRabbit UI

Review profile: CHILL

Plan: Pro

📥 Commits

Reviewing files that changed from the base of the PR and between 2c80f03 and 712759a.

📒 Files selected for processing (4)
  • src/Interpretability/DemographicParityBiasDetector.cs (1 hunks)
  • src/Interpretability/DisparateImpactBiasDetector.cs (1 hunks)
  • src/Interpretability/EqualOpportunityBiasDetector.cs (1 hunks)
  • src/Models/Results/PredictionModelResult.cs (4 hunks)
🚧 Files skipped from review as they are similar to previous changes (1)
  • src/Interpretability/EqualOpportunityBiasDetector.cs
🧰 Additional context used
🧬 Code graph analysis (2)
src/Interpretability/DisparateImpactBiasDetector.cs (1)
src/Interpretability/InterpretabilityMetricsHelper.cs (1)
  • InterpretabilityMetricsHelper (25-288)
src/Interpretability/DemographicParityBiasDetector.cs (1)
src/Interpretability/InterpretabilityMetricsHelper.cs (1)
  • InterpretabilityMetricsHelper (25-288)
⏰ Context from checks skipped due to timeout of 90000ms. You can increase the timeout in your CodeRabbit configuration to a maximum of 15 minutes (900000ms). (1)
  • GitHub Check: CodeQL analysis (csharp)
🔇 Additional comments (3)
src/Models/Results/PredictionModelResult.cs (3)

166-176: LGTM! Properties follow existing patterns.

The new BiasDetector and FairnessEvaluator properties are properly defined as nullable with private setters, consistent with the class's design.


208-219: LGTM! Constructor properly initializes new properties.

The optional parameters with default null values and corresponding property assignments are correctly implemented.


427-428: LGTM! Deserialization properly restores new properties.

The properties are correctly restored during deserialization, maintaining consistency with other properties.

Comment on lines +34 to +38
protected override BiasDetectionResult<T> GetBiasDetectionResult(
Vector<T> predictions,
Vector<T> sensitiveFeature,
Vector<T>? actualLabels)
{

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🔴 Critical

Add input validation for vector lengths.

The method lacks validation that predictions and sensitiveFeature have matching lengths. When GetSubset is called at line 58, it uses indices derived from sensitiveFeature to access elements in predictions, which will cause an IndexOutOfRangeException if the vectors have different lengths.

Apply this diff to add validation at the start of the method:

 protected override BiasDetectionResult<T> GetBiasDetectionResult(
     Vector<T> predictions,
     Vector<T> sensitiveFeature,
     Vector<T>? actualLabels)
 {
+    if (predictions.Length != sensitiveFeature.Length)
+    {
+        throw new ArgumentException(
+            $"Predictions length ({predictions.Length}) must match sensitiveFeature length ({sensitiveFeature.Length})");
+    }
+
     var result = new BiasDetectionResult<T>();
📝 Committable suggestion

‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.

Suggested change
protected override BiasDetectionResult<T> GetBiasDetectionResult(
Vector<T> predictions,
Vector<T> sensitiveFeature,
Vector<T>? actualLabels)
{
protected override BiasDetectionResult<T> GetBiasDetectionResult(
Vector<T> predictions,
Vector<T> sensitiveFeature,
Vector<T>? actualLabels)
{
if (predictions.Length != sensitiveFeature.Length)
{
throw new ArgumentException(
$"Predictions length ({predictions.Length}) must match sensitiveFeature length ({sensitiveFeature.Length})");
}
var result = new BiasDetectionResult<T>();
🤖 Prompt for AI Agents
In src/Interpretability/DemographicParityBiasDetector.cs around lines 34 to 38,
the method lacks validation that the input vectors have matching lengths which
can cause IndexOutOfRangeException when GetSubset is called; add checks at the
start of the method to validate predictions and sensitiveFeature are not null
and that predictions.Length == sensitiveFeature.Length (and if actualLabels is
not null, assert actualLabels.Length == predictions.Length), and if any check
fails throw an ArgumentException (or ArgumentNullException for null inputs) with
a clear message indicating which vectors have mismatched lengths.

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull Request Overview

Copilot reviewed 17 out of 18 changed files in this pull request and generated 6 comments.


💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.

Comment on lines +82 to +85
else
{
result.DisparateImpactRatio = _numOps.Divide(minRate, maxRate);
}

Copilot AI Nov 1, 2025

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The 'else' keyword is redundant here because the 'if' block at line 75 contains a return statement. The code after the if block will only execute when the condition is false, making the 'else' unnecessary. Remove the 'else' and unindent the code block for cleaner, more maintainable code.

Suggested change
else
{
result.DisparateImpactRatio = _numOps.Divide(minRate, maxRate);
}
result.DisparateImpactRatio = _numOps.Divide(minRate, maxRate);

Copilot uses AI. Check for mistakes.
T minRate = orderedRates.First();
T maxRate = orderedRates.Last();

// Avoid division by zero

Copilot AI Nov 1, 2025

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The comment on line 74 says 'Avoid division by zero', but the check at line 75 also handles the edge case where all groups have zero positive predictions (no one receives positive outcomes). Consider updating the comment to 'Handle edge case where all predictions are zero' to more accurately describe what this condition represents, as it's not just about avoiding division by zero but also about the semantic meaning of having no positive predictions across all groups.

Suggested change
// Avoid division by zero
// Handle edge case where all predictions are zero

Copilot uses AI. Check for mistakes.
Comment on lines +103 to +108
foreach (var group in groups)
{
string groupKey = group?.ToString() ?? "unknown";
fairnessMetrics.AdditionalMetrics[$"Group_{groupKey}_PositiveRate"] = groupPositiveRates[groupKey];
fairnessMetrics.AdditionalMetrics[$"Group_{groupKey}_Size"] = _numOps.FromDouble(groupSizes[groupKey]);
}

Copilot AI Nov 1, 2025

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This foreach loop immediately maps its iteration variable to another variable - consider mapping the sequence explicitly using '.Select(...)'.

Suggested change
foreach (var group in groups)
{
string groupKey = group?.ToString() ?? "unknown";
fairnessMetrics.AdditionalMetrics[$"Group_{groupKey}_PositiveRate"] = groupPositiveRates[groupKey];
fairnessMetrics.AdditionalMetrics[$"Group_{groupKey}_Size"] = _numOps.FromDouble(groupSizes[groupKey]);
}
var groupPositiveRateMetrics = groups
.Select(group => group?.ToString() ?? "unknown")
.ToDictionary(
groupKey => $"Group_{groupKey}_PositiveRate",
groupKey => groupPositiveRates[groupKey]
);
var groupSizeMetrics = groups
.Select(group => group?.ToString() ?? "unknown")
.ToDictionary(
groupKey => $"Group_{groupKey}_Size",
groupKey => _numOps.FromDouble(groupSizes[groupKey])
);
foreach (var kvp in groupPositiveRateMetrics)
{
fairnessMetrics.AdditionalMetrics[kvp.Key] = kvp.Value;
}
foreach (var kvp in groupSizeMetrics)
{
fairnessMetrics.AdditionalMetrics[kvp.Key] = kvp.Value;
}

Copilot uses AI. Check for mistakes.
Comment on lines +80 to +87
if (result.HasBias)
{
result.Message = $"Bias detected: Statistical parity difference = {statisticalParityValue:F3} (exceeds {_threshold} threshold)";
}
else
{
result.Message = $"No significant bias detected: Statistical parity difference = {statisticalParityValue:F3}";
}

Copilot AI Nov 1, 2025

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Both branches of this 'if' statement write to the same variable - consider using '?' to express intent better.

Suggested change
if (result.HasBias)
{
result.Message = $"Bias detected: Statistical parity difference = {statisticalParityValue:F3} (exceeds {_threshold} threshold)";
}
else
{
result.Message = $"No significant bias detected: Statistical parity difference = {statisticalParityValue:F3}";
}
result.Message = result.HasBias
? $"Bias detected: Statistical parity difference = {statisticalParityValue:F3} (exceeds {_threshold} threshold)"
: $"No significant bias detected: Statistical parity difference = {statisticalParityValue:F3}";

Copilot uses AI. Check for mistakes.
Comment on lines +91 to +98
if (result.HasBias)
{
result.Message = $"Bias detected: Disparate impact ratio = {disparateImpactValue:F3} (below {_threshold} threshold)";
}
else
{
result.Message = $"No significant bias detected: Disparate impact ratio = {disparateImpactValue:F3}";
}

Copilot AI Nov 1, 2025

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Both branches of this 'if' statement write to the same variable - consider using '?' to express intent better.

Suggested change
if (result.HasBias)
{
result.Message = $"Bias detected: Disparate impact ratio = {disparateImpactValue:F3} (below {_threshold} threshold)";
}
else
{
result.Message = $"No significant bias detected: Disparate impact ratio = {disparateImpactValue:F3}";
}
result.Message = result.HasBias
? $"Bias detected: Disparate impact ratio = {disparateImpactValue:F3} (below {_threshold} threshold)"
: $"No significant bias detected: Disparate impact ratio = {disparateImpactValue:F3}";

Copilot uses AI. Check for mistakes.
Comment on lines +99 to +106
if (result.HasBias)
{
result.Message = $"Bias detected: Equal opportunity difference = {eoValue:F3} (exceeds {_threshold} threshold)";
}
else
{
result.Message = $"No significant bias detected: Equal opportunity difference = {eoValue:F3}";
}

Copilot AI Nov 1, 2025

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Both branches of this 'if' statement write to the same variable - consider using '?' to express intent better.

Suggested change
if (result.HasBias)
{
result.Message = $"Bias detected: Equal opportunity difference = {eoValue:F3} (exceeds {_threshold} threshold)";
}
else
{
result.Message = $"No significant bias detected: Equal opportunity difference = {eoValue:F3}";
}
result.Message = result.HasBias
? $"Bias detected: Equal opportunity difference = {eoValue:F3} (exceeds {_threshold} threshold)"
: $"No significant bias detected: Equal opportunity difference = {eoValue:F3}";

Copilot uses AI. Check for mistakes.
Add input validation to ensure predictions and sensitiveFeature vectors
have matching lengths before processing. Prevents runtime errors from
mismatched vector access.

Changes:
- DisparateImpactBiasDetector: validate vector lengths match
- DemographicParityBiasDetector: validate vector lengths match
- EqualOpportunityBiasDetector: validate vector lengths match

Uses optimized Length property instead of Count() for performance.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude <noreply@anthropic.com>
@ooples
ooples merged commit 9b315c8 into master Nov 1, 2025
3 of 5 checks passed
@ooples
ooples deleted the feat/us-nf-011-ethical-ai-tools branch November 1, 2025 04:59
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants