Multi-stage evidence-grounded inference architecture replacing the monolithic 9B model extraction pipeline. CPU-first specialist services handle routine extraction while the 9B vLLM model is preserved for semantic adjudication of ambiguous cases. Key components: - Capability-aware inference gateway (OpenAI-compatible + Ollama) - Endpoint registry with DB migrations and REST API - Sentence-aware document segmenter (property tests) - Deterministic financial parsing with offset integrity - Symbol resolution with ambiguity detection - Specialist service (GLiNER2, dynamic batching, K8s deployment) - Company-specific sentiment (FinBERT, calibration) - Retrieval-based novelty and duplicate detection - Confidence calibration pipeline - Deterministic routing engine (property tests) - 9B adjudication layer with VRAM gating - Stock-specific impact model (features, labels, baseline, trained) - Pipeline orchestrator (state machine, queues, leases, feature flags) - Bounded parallelism (async workers, semaphore, load shedding) - Observability (tracing, metrics, alerts) - Compatibility adapter (v3→v2 golden mapping tests) - Shadow/canary promotion framework - Active learning and fine-tuning pipeline Test results: 1,161 tests pass, ruff lint clean. All 282 spec tasks completed.
14 KiB
V3 Annotation Guidelines
Schema version: 1.0.0 Last updated: 2025-01-15
Purpose
These guidelines define how human annotators and automated systems label documents in the Intelligence Pipeline v3 Gold Corpus. Every annotation must be evidence-grounded — no label is valid without a supporting evidence span traceable to the source text.
Core Principles
- Evidence first. If you cannot point to exact text that supports a label, do not apply the label.
- Explicit over inferred. Mark only what the document explicitly states in primary annotations. Inferred exposure uses a separate, lower-confidence channel.
- Precision over recall. A missed entity is preferable to a fabricated one. The pipeline uses multiple stages — later stages catch omissions.
- Reproducibility. Two annotators given the same document should produce substantially the same labels. Ambiguous cases are marked, not resolved by guess.
Evidence Spans
Definition
An evidence span is the exact substring of the source document that supports an annotation. It uses zero-based character offsets into the original (pre-chunking) document text.
Rules
- Every entity, event, relation, numeric fact, and sentiment annotation MUST reference at least one evidence span.
- Spans should be minimal but complete — include enough context for the label to be verifiable without the full document.
- Overlapping spans are permitted (e.g., the same sentence supports both an entity and an event).
- The
textfield MUST exactly matchsource_text[start_char:end_char].
Positive example
Source: "Apple Inc. reported quarterly earnings of $1.52 per share"
Span: start_char=0, end_char=10, text="Apple Inc."
Negative example
Source: "Apple Inc. reported quarterly earnings of $1.52 per share"
Span: start_char=0, end_char=5, text="Apple"
❌ Truncating "Apple Inc." to "Apple" loses the corporate suffix needed to distinguish from Apple Records or the fruit.
Entity Annotation
Entity Types
| Type | When to use | Example |
|---|---|---|
company |
Legal entity, publicly traded firm, government agency | "Apple Inc.", "The Federal Reserve" |
person |
Named individual | "Tim Cook", "Jerome Powell" |
product |
Named product or service | "iPhone 16", "Azure OpenAI Service" |
event |
Named event instance | "Q1 2025 earnings call" |
financial_metric |
Named metric class | "EPS", "revenue", "free cash flow" |
date |
Temporal expression | "Q1 2025", "January 15, 2025" |
percentage |
Percentage value | "4%", "25 basis points" |
currency |
Monetary value | "$1.52", "$10 billion" |
relationship |
Explicit relationship mention | "subsidiary", "joint venture partner" |
Canonical Resolution
- If an entity maps to a company in the symbol registry, set
canonical_idandcanonical_name(ticker). - If an entity is ambiguous (e.g., "Apple" could be AAPL or a fruit company), mark an ambiguity marker and set confidence below 1.0.
- Do NOT invent canonical IDs. If not in the registry, leave
canonical_idas null.
Positive example
{
"entity_type": "company",
"literal_text": "Alphabet",
"canonical_id": "googl-uuid",
"canonical_name": "GOOGL",
"confidence": 0.97
}
Negative example
{
"entity_type": "company",
"literal_text": "the company",
"canonical_id": "aapl-uuid",
"canonical_name": "AAPL",
"confidence": 0.90
}
❌ "the company" is a pronoun reference, not an entity mention. Resolve coreference but annotate the actual named mention, not the pronoun.
Event Classification
Event Classes
| Class | Definition | Distinguishing criteria |
|---|---|---|
earnings_beat |
Reported EPS or revenue exceeds consensus | Explicit comparison to estimates |
earnings_miss |
Reported EPS or revenue below consensus | Explicit comparison to estimates |
guidance_raise |
Forward guidance raised vs prior or consensus | Future-looking, not historical result |
guidance_cut |
Forward guidance lowered | Future-looking, not historical result |
ma_announcement |
Merger, acquisition, investment, or divestiture | Transaction between entities |
legal_regulatory |
Lawsuit, fine, regulatory action, or settlement | Legal or regulatory body involved |
product_launch |
New product, service, or major feature announced | Not routine updates |
supply_chain |
Disruption, partnership, or change in supply relationships | Affects production/delivery |
rating_change |
Analyst upgrade, downgrade, or target change | From research analyst/firm |
management_change |
CEO/CFO/board appointment, resignation, or removal | C-suite or board level |
macro_event |
Interest rates, policy, trade, geopolitical | Not specific to one company |
dividend_change |
Dividend increase, decrease, or special dividend | Shareholder distribution |
buyback |
Share repurchase program announcement or completion | Capital return via buyback |
Adjudication triggers for events
Route to the 9B adjudicator when:
- The same facts could be classified as multiple event types (e.g., guidance_raise during an earnings call could be either earnings_beat or guidance_raise — label the most specific applicable class).
- The event is implied but not explicitly stated.
- The primary company is unclear.
Positive example
Source: "Apple beat earnings expectations with EPS of $1.52 vs $1.43 expected"
Event class: earnings_beat
Confidence: 0.98
Negative example
Source: "Apple reported EPS of $1.52"
Event class: earnings_beat
❌ Without a comparison to consensus/estimates, this is a numeric fact report, not an earnings beat. The document must provide evidence of beating expectations.
Relations
Relation Types
| Type | Subject | Object | When to use |
|---|---|---|---|
directly_affects |
Event | Company | Event explicitly names or discusses the company |
inferred_exposure |
Event | Company | Exposure inferred from sector, supply chain, or competition |
competes_with |
Company | Company | Competitive relationship stated or clearly implied |
supplies |
Company | Company | Supply chain relationship stated |
Critical distinction: directly_affects vs inferred_exposure
directly_affects: The document explicitly states the company is impacted. Evidence span exists.inferred_exposure: The impact is reasoned from relationships, not stated. May have weak or no direct evidence span.
Only directly_affects enters primary company extraction. inferred_exposure flows through the separate interpolation/propagation architecture with distinct confidence and provenance.
Positive example (directly_affects)
Source: "Microsoft announced a $10 billion investment in OpenAI"
Relation: directly_affects(event=ma_announcement, company=Microsoft)
Evidence: "Microsoft announced"
Negative example (incorrectly using directly_affects)
Source: "Microsoft announced a $10 billion investment in OpenAI"
Relation: directly_affects(event=ma_announcement, company=Google)
❌ Google is not mentioned in the event sentence. This should be inferred_exposure based on competitive relationship, with appropriate lower confidence.
Numeric Facts
Annotation rules
- Always store both
literal_value(exact text) andnormalized_value(parsed number). - Include
unit(USD, %, bps, shares, etc.). - Link to the subject entity when determinable.
- Use
predicateto capture the semantic role: reported, expected, raised_to, cut_to, beat_by, missed_by. - Include
periodwhen the fact references a specific time frame.
Normalization conventions
| Literal | Normalized | Unit |
|---|---|---|
| "$1.52" | 1.52 | USD |
| "$94.9 billion" | 94900000000 | USD |
| "25 basis points" | 0.25 | percentage_points |
| "4%" | 4.0 | % |
| "$0.26 per share" | 0.26 | USD |
Positive example
{
"fact_type": "eps",
"predicate": "reported",
"literal_value": "$1.52 per share",
"normalized_value": 1.52,
"unit": "USD",
"period": {"period_type": "fiscal_quarter", "fiscal_year": 2025, "fiscal_quarter": 1}
}
Negative example
{
"fact_type": "eps",
"predicate": "reported",
"literal_value": "$1.52 per share",
"normalized_value": 152,
"unit": "cents"
}
❌ While $1.52 = 152 cents, always normalize to the unit stated in the source. Conversion to a different unit introduces potential confusion.
Sentiment
Rules
- Sentiment is company-specific, not document-level. A single article can have positive sentiment for one company and negative for another.
- Annotate probability distributions (positive, negative, neutral) that sum to 1.0.
mixedlabel is used when evidence groups disagree — it is computed from evidence-group-level disagreement, NOT an unconstrained fourth class.- The label should reflect the dominant probability.
When to label "mixed"
Label mixed when:
- Different paragraphs contain opposing sentiment for the same company
- The same fact has both positive and negative implications (e.g., restructuring = cost cuts but also layoffs)
- Analyst opinions explicitly disagree within the document
Do NOT label mixed when:
- Sentiment is merely uncertain or mild — that's
neutralwith lower confidence - The document discusses multiple companies with different sentiments — annotate separately per company
Positive example
{
"label": "mixed",
"positive_probability": 0.40,
"negative_probability": 0.45,
"neutral_probability": 0.15,
"evidence_ids": ["ev-pressure", "ev-validation"]
}
(Article says AI investment pressures cloud revenue but validates the broader thesis)
Negative example
{
"label": "mixed",
"positive_probability": 0.85,
"negative_probability": 0.05,
"neutral_probability": 0.10
}
❌ When positive_probability dominates at 0.85, the label should be positive, not mixed. Mixed requires genuine disagreement in evidence.
Direct Effects vs Inferred Exposure
Direct Effects
A direct effect means the document explicitly states or clearly demonstrates that an event impacts a specific company.
Criteria:
- The company is named in the same sentence or paragraph as the event
- The causal link is stated, not inferred
- Evidence span directly connects event to company
Inferred Exposure
Inferred exposure captures reasoned but unstated impacts on companies.
Criteria:
- The company is NOT explicitly linked to the event in the source text
- The connection comes from known relationships (competitor, supplier, sector peer)
- Confidence should be lower than direct effects (typically 0.5–0.8)
- Requires
reasoningfield explaining the inference chain
Adjudication routing
When it's unclear whether an effect is direct or inferred, mark an ambiguity marker with type implied_causal_impact and route to the 9B adjudicator.
Ambiguity Markers
When to flag
Flag ambiguity when:
- An alias resolves to multiple candidate companies (
unresolved_alias) - Multiple companies could be the primary subject (
multiple_primary_companies) - Numeric facts within the same document contradict each other (
contradictory_numeric_facts) - Sentiment evidence points in opposing directions for the same company (
conflicting_sentiment) - Impact is implied through causal chain, not stated (
implied_causal_impact) - Guidance must be compared to consensus to determine direction (
guidance_vs_consensus_requires_reasoning) - A required field cannot be determined from available evidence (
material_field_missing) - Evidence covers less than the minimum threshold for confident extraction (
evidence_coverage_below_threshold) - Calibrated confidence falls below the routing threshold (
calibrated_confidence_below_threshold) - A relation spans multiple document chunks (
long_document_cross_chunk_relation)
Severity levels
- low: The annotation is likely correct but has reduced certainty. Fast path may proceed with a confidence penalty.
- medium: The annotation requires review. Routes to adjudication by default.
- high: The annotation cannot be reliably made without semantic reasoning. Always routes to adjudication.
Safety-Critical Fields
The following fields are safety-critical for promotion gates. Errors in these fields can directly cause incorrect trading decisions:
| Field | Why it's critical | Minimum promotion gate |
|---|---|---|
| Company identity (ticker) | Wrong ticker = trade on wrong security | Precision ≥ 0.95, Recall ≥ 0.90 |
| Event class | Misclassifying beat/miss inverts signal direction | Macro-F1 ≥ 0.85 |
| Sentiment direction | Wrong sentiment → wrong position direction | Direction accuracy ≥ 0.90 |
| Numeric fact values | Wrong magnitude affects impact estimation | Tolerance match ≥ 0.92 |
| Direct effect attribution | Wrong company attribution creates false signals | Precision ≥ 0.93 |
| Evidence support | Unsupported claims are unverifiable | Support rate ≥ 0.95 |
| Confidence calibration | Overconfidence bypasses review | ECE ≤ 0.05 |
Annotators must pay special attention to these fields. During review, any error in a safety-critical field requires correction before the annotation can receive "gold" status.
Annotation Workflow
- First pass: Identify all entities and evidence spans
- Second pass: Classify events and link to companies
- Third pass: Extract numeric facts with periods
- Fourth pass: Assess per-company sentiment
- Fifth pass: Identify relations, direct effects, and inferred exposures
- Sixth pass: Flag ambiguities and set confidence levels
- Review: Senior annotator validates safety-critical fields
Inter-annotator agreement
Hard cases (flagged with ambiguity markers) receive double annotation. Inter-annotator agreement is measured per field type using Cohen's kappa. Target: κ ≥ 0.80 for entity and event labels, κ ≥ 0.70 for relations and sentiment.
Version History
| Version | Date | Changes |
|---|---|---|
| 1.0.0 | 2025-01-15 | Initial schema and guidelines |