FinRobot Pro
GitHub Sign in

FinRobot · Agent Architecture · 2026

How we built it

From a model that writes the numbers,
to a model that can only call for numbers already computed.

FinRobot runs two agent stacks side by side. V1 is a workflow with model calls in it; V2 is an agent harness. That is the difference everything else on this page follows from — including why V2 had to take away the model's ability to produce a number. Every claim below can be checked against the source.

01 · The distinction that matters

Workflow, or harness?

The word “agent” covers two very different things, and what separates them is who decides what happens next. In a workflow the code decides: the author fixes the sequence and the model fills in slots. In a harness the model decides: the code supplies tools, guardrails and a loop, and the model picks its next move each turn. V1 is the first. V2 is the second.

V1 · a workflow with model calls in it code assembles LLM ×8 fills slots code renders Control never leaves the code. A new capability means editing the pipeline. V2 · a harness the model drives model tool or a whole pipeline chooses observes The model decides the next step, every turn. A new capability means adding a tool.
The loop on the right is the harness. Not the model, and not the prompt — the machinery that lets a model act, see what happened, and act again. V1 has no such loop: its eight calls are made by code, and none of them ever learns what the others produced.
Workflow · V1 Harness · V2
Who picks the next step the author, when the code was written the model, on each turn of the loop
Does the model see results No. It returns its slot and is done Yes — a tool result comes back as an observation
Recovering from bad output the step fails ModelRetry hands back the reason and it tries again
Adding a capability edit the pipeline add a tool
Where a fixed sequence lives it is the program it is one tool among fifteen

The pipeline did not go away — it became a tool. This is the part worth being precise about, because “harness” is not a synonym for open-ended improvisation. V2's eight-step research pipeline is still a fixed sequence, deliberately: a report people rely on should not be assembled a different way each time. What changed is where that sequence sits. In V1 it was the program. In V2 it is one of seven pipelines, each generated from the pipeline registry as a tool the agent can choose to call — alongside eight conversational tools and, when nothing fits, a skill it can load and improvise within.

And the harness is why the numbers had to be walled off. These two ideas are connected, not parallel. Once the model picks its own path you can no longer enumerate what it will do — so the facts have to be put somewhere it cannot reach. Handing the model control over the process is what forced taking away its ability to produce a number. That is what the compute boundary further down is for: a harness without hard walls around its facts is just a faster way to be confidently wrong.

02 · The stacks

Two agent stacks, two assumptions about where numbers come from

Framework choice is not a matter of taste. The two versions hand “who produces the number” to different things — V1 hands it to a model, V2 hands it to a function.

V1

Report Studio

Framework
openai-agents (the OpenAI Agents SDK)
SDK pin
requires openai<2
Agents
8, all peers
Typed output
every agent declares an output_type, so it returns a validated Pydantic model rather than free text
Tool calls
none. tools= and handoffs= appear in none of the 8 agents
How data arrives
Python flattens the financials into one text prompt first
Front end
Jinja2 + Vue 3 (CDN)

V2

Research Desk

Framework
pydantic-ai >= 1.70
SDK pin
requires openai >= 2.25
Agents
one lead agent + 5 role sub-agents: data · analysis · modeling · synthesis · report
Type system
Agent[FinRobotDeps, str] — dependencies injected through the generic, tool signatures type-checked
Tool calls
15 tools on the lead agent: 8 conversational, plus 7 pipeline dispatchers generated from the pipeline registry
How data arrives
26 deterministic operators, the only way a model obtains a number
Front end
React + Vite (plus a Tauri desktop shell)

03 · V1 architecture

Eight independent calls, one shared prompt

V1 was a reasonable design in 2025. It got typed output right — making a model return validated objects was not yet standard practice. Its ceiling is elsewhere: an agent cannot ask for anything, so what is not in the prompt does not exist to it.

one prompt · eight independent calls FMP financials flattened into one text prompt tagline company_overview investment_overview valuation_overview risks competitor_analysis major_takeaways news_summary typed fragments fixed template report PDF
Not one arrow points back from an agent to a data source. Each of the eight runs once through Runner.run(agent, prompt) — no shared state, no follow-up question, no way to fetch what is missing. Every number in the valuation section was produced by a model doing arithmetic in tokens.
  • What it got rightoutput_type makes each agent return a validated Pydantic object, instead of free text somebody has to parse with regexes.
  • The ceilingData is delivered once, up front. Wanting a specific sentence out of the 10-K, or a different discount-rate assumption, leads nowhere.
  • The consequenceThe report is fixed at 13 chapters, because the pipeline is one-directional — it cannot decide what to do next from what it just saw.

04 · V2 architecture

Layered, with a tool boundary

V2 separates deciding what to do from doing it. The lead agent only orchestrates and judges; five role sub-agents carry the steps; and every number has to cross the tool boundary and be produced by a deterministic operator. A model cannot write a valuation out of thin air.

question / ticker Lead Agent 15 tools · orchestration & judgment data analysis modeling synthesis report fetch peers · catalysts DCF · technicals thesis write-up compute boundary Deterministic compute engine · 26 operators dcf · ddm · lbo · wacc · monte_carlo · multiples sotp · residual_income · xbrl_aligned_comps · … Data layer FMP · yfinance · SEC EDGAR cross-checked · divergence flagged call an operator fetch raw data tool boundary
The two dashed lines are the whole point of this diagram. The upper one separates orchestration from execution; the lower one fixes where numbers come from — nothing above it can produce one. Amber is written by a model, cyan is computed by code, and the two do not overlap.
  • 26 deterministic operatorsDCF, DDM, LBO, WACC, Monte Carlo, comparables, sum-of-the-parts, residual income — each a pure function with unit tests.
  • 8 conversational toolsquery_financial_data · ask_filings · run_monte_carlo · run_backtest · find_reports · diff_reports · query_coverage_universe · activate_skill
  • 7 pipeline toolsGenerated from the pipeline registry, so a whole analysis is something the agent calls: run_equity_research · run_dcf_valuation · run_lbo_analysis · run_ddm_valuation · run_comps_analysis · run_earnings_analysis · run_ic_memo
  • ModelRetryWhen validation fails, the reason goes back to the model and it tries again — instead of failing the run or passing bad data downstream.
  • 7 domain skill packsEquity research, financial analysis, investment banking, private equity, wealth management, plus two vendor packs — loaded on demand by activate_skill.

05 · The pipeline

The eight steps behind one report

This is the most concrete layer of “how we built it”. Each step is bound to one role agent, produces a validated structured object, and passes it forward as context — which is why a failure lands on a named step rather than on “report generation failed”. This is one of seven pipelines, and the lead agent reaches it the same way it reaches anything else: by calling run_equity_research.

# Step Role What it does
1 data_collection data Pull filings, prices and statements; cross-check the sources, emit structured financials
2 catalyst_analysis data Identify earnings dates, guidance changes and other event drivers
3 peer_analysis analysis Screen peers and normalise comparables (xbrl_aligned_comps)
4 financial_modeling modeling DCF / DDM / LBO / comparables — all computed by operators, then cross-reconciled
5 ownership_governance_analysis analysis 13F institutional holdings, Form 4 insider trades, governance structure
6 technical_analysis modeling Price momentum, volatility and technical signals
7 thesis synthesis Combine all of the above into an investment thesis and a rating
8 report report Write up: a 13-chapter report in which every number carries its origin

The operators argue with each other. At step 4, a method whose result strays too far from the cross-method median is flagged as an outlier rather than quietly averaged in. From a real production log: ev_ebitda mid $261.20 deviates 42% from cross-method median $184.43 — flagged as outlier. That check lives in code; it does not depend on a model noticing.

06 · Why the design matters

Every number traces back to what produced it

In financial research, “where did this number come from” is the audit floor. A fair value written by an LLM cannot answer it — there are no intermediate steps, only tokens. In V2 every number has a retraceable operator chain, and the model's job is what sits outside that chain.

SEC EDGAR 10-K / 10-Q xbrl_aligned normalise wacc discount rate dcf discount flows fair value $182.40 every step a unit-testable pure function · inputs and versions recorded in provenance numbers / judgment synthesis agent writes only the why Thesis · risks · catalysts cites numbers, never produces them as input
The model has no arrow into the chain above. It receives $182.40 already computed; its task is to explain what that number means, not to produce it. Which is why “the model got the arithmetic wrong” is not a failure mode that exists here.
  • AuditableEvery number traces to the operator, the inputs and the source version that produced it. In a compliance setting that is a requirement, not a bonus.
  • ReproducibleSame inputs, same operators, same number. The model's temperature cannot reach it.
  • Cross-checkedThe same field is fetched from both FMP and yfinance and flagged when they diverge past a threshold — production has logged a cash figure differing by 37%.
  • InteractiveBecause the DCF is a function, a user can change an assumption and recompute. V1 cannot; its valuation section is prose that has already been written.

07 · The comparison

Four differences you can verify

“Better” is not worth claiming if it only means it feels better. Each row below corresponds to a structural difference.

V1 · Report Studio V2 · Research Desk Why it is a difference
Where numbers originate a model doing arithmetic in tokens 26 deterministic operators Arithmetic errors and invented figures are structurally off the number path
Can it ask for more No — nothing exists outside the prompt 15 tools: query filings, run a Monte Carlo, diff two saved reports, dispatch a whole pipeline The information that decides the next step can be obtained during the run, not staged in advance
How it is organised 8 peer agents, none aware of the others lead + 5 roles, an 8-step pipeline, each step bound to a role Prompts can be scoped per role, and a failure localises to one step
Failure handling validation fails → the step fails ModelRetry returns the reason and the model retries Bad data does not slip silently downstream — the most common failure in agent systems
What comes out a fixed 13-chapter PDF a traceable desk where assumptions can be changed and recomputed Research is iterative; a one-shot document does not serve it

Honestly, V2 is worse in places. It is slower and more expensive — layering plus tool calls means more model round-trips. Its dependency tree is heavier, and a cold start has to warm the compute engine and a symbol index. And for “just give me a standard report”, V1's one-directional pipeline is still faster, cheaper and more predictable. So the two coexist as two ways of working, not as an old version and its replacement — which is why this site keeps both entrances.