Files
Celes Renata a72f336ad1 feat: Intelligence Pipeline v3 — full implementation
Multi-stage evidence-grounded inference architecture replacing the
monolithic 9B model extraction pipeline. CPU-first specialist services
handle routine extraction while the 9B vLLM model is preserved for
semantic adjudication of ambiguous cases.

Key components:
- Capability-aware inference gateway (OpenAI-compatible + Ollama)
- Endpoint registry with DB migrations and REST API
- Sentence-aware document segmenter (property tests)
- Deterministic financial parsing with offset integrity
- Symbol resolution with ambiguity detection
- Specialist service (GLiNER2, dynamic batching, K8s deployment)
- Company-specific sentiment (FinBERT, calibration)
- Retrieval-based novelty and duplicate detection
- Confidence calibration pipeline
- Deterministic routing engine (property tests)
- 9B adjudication layer with VRAM gating
- Stock-specific impact model (features, labels, baseline, trained)
- Pipeline orchestrator (state machine, queues, leases, feature flags)
- Bounded parallelism (async workers, semaphore, load shedding)
- Observability (tracing, metrics, alerts)
- Compatibility adapter (v3→v2 golden mapping tests)
- Shadow/canary promotion framework
- Active learning and fine-tuning pipeline

Test results: 1,161 tests pass, ruff lint clean.
All 282 spec tasks completed.
2026-07-13 02:14:59 +00:00

47 lines
1.2 KiB
Python

"""Benchmark configuration and comparison framework for Intelligence Pipeline v3.
Defines extraction configurations for controlled comparison between the current
production pipeline and corrected variants. Supports attribution of improvement
sources (temperature fix, schema constraints, architecture changes).
Validates: Requirements 16.2, 16.3, 16.5
"""
from services.intelligence_pipeline_v3.benchmark.comparison import (
ComparisonReport,
ConfigDelta,
FieldDelta,
ResourceDelta,
compare_configurations,
)
from services.intelligence_pipeline_v3.benchmark.configurations import (
BASELINE_CURRENT,
BASELINE_STRICT_SCHEMA,
BASELINE_TEMP_ZERO,
BenchmarkConfig,
StructuredOutputMode,
list_configurations,
)
from services.intelligence_pipeline_v3.benchmark.runner import (
BenchmarkDocumentResult,
BenchmarkRun,
BenchmarkRunner,
)
__all__ = [
"BASELINE_CURRENT",
"BASELINE_STRICT_SCHEMA",
"BASELINE_TEMP_ZERO",
"BenchmarkConfig",
"BenchmarkDocumentResult",
"BenchmarkRun",
"BenchmarkRunner",
"ComparisonReport",
"ConfigDelta",
"FieldDelta",
"ResourceDelta",
"StructuredOutputMode",
"compare_configurations",
"list_configurations",
]