FinRobot · Agent Architecture · 2026
How we built it
From a model that writes the numbers,
to a model that can only call for numbers already computed.
FinRobot runs two agent stacks side by side. V1 is a workflow with model calls in it; V2 is an agent harness. That is the difference everything else on this page follows from — including why V2 had to take away the model's ability to produce a number. Every claim below can be checked against the source.
01 · The distinction that matters
Workflow, or harness?
The word “agent” covers two very different things, and what separates them is who decides what happens next. In a workflow the code decides: the author fixes the sequence and the model fills in slots. In a harness the model decides: the code supplies tools, guardrails and a loop, and the model picks its next move each turn. V1 is the first. V2 is the second.
| Workflow · V1 | Harness · V2 | |
|---|---|---|
| Who picks the next step | the author, when the code was written | the model, on each turn of the loop |
| Does the model see results | No. It returns its slot and is done | Yes — a tool result comes back as an observation |
| Recovering from bad output | the step fails | ModelRetry hands back the reason and it tries again |
| Adding a capability | edit the pipeline | add a tool |
| Where a fixed sequence lives | it is the program | it is one tool among fifteen |
The pipeline did not go away — it became a tool. This is the part worth being precise about, because “harness” is not a synonym for open-ended improvisation. V2's eight-step research pipeline is still a fixed sequence, deliberately: a report people rely on should not be assembled a different way each time. What changed is where that sequence sits. In V1 it was the program. In V2 it is one of seven pipelines, each generated from the pipeline registry as a tool the agent can choose to call — alongside eight conversational tools and, when nothing fits, a skill it can load and improvise within.
And the harness is why the numbers had to be walled off. These two ideas are connected, not parallel. Once the model picks its own path you can no longer enumerate what it will do — so the facts have to be put somewhere it cannot reach. Handing the model control over the process is what forced taking away its ability to produce a number. That is what the compute boundary further down is for: a harness without hard walls around its facts is just a faster way to be confidently wrong.
02 · The stacks
Two agent stacks, two assumptions about where numbers come from
Framework choice is not a matter of taste. The two versions hand “who produces the number” to different things — V1 hands it to a model, V2 hands it to a function.
V1
Report Studio
- Framework
openai-agents(the OpenAI Agents SDK)- SDK pin
- requires
openai<2 - Agents
- 8, all peers
- Typed output
- every agent declares an
output_type, so it returns a validated Pydantic model rather than free text - Tool calls
- none.
tools=andhandoffs=appear in none of the 8 agents - How data arrives
- Python flattens the financials into one text prompt first
- Front end
- Jinja2 + Vue 3 (CDN)
V2
Research Desk
- Framework
pydantic-ai >= 1.70- SDK pin
- requires
openai >= 2.25 - Agents
- one lead agent + 5 role sub-agents: data · analysis · modeling · synthesis · report
- Type system
Agent[FinRobotDeps, str]— dependencies injected through the generic, tool signatures type-checked- Tool calls
- 15 tools on the lead agent: 8 conversational, plus 7 pipeline dispatchers generated from the pipeline registry
- How data arrives
- 26 deterministic operators, the only way a model obtains a number
- Front end
- React + Vite (plus a Tauri desktop shell)
03 · V1 architecture
Eight independent calls, one shared prompt
V1 was a reasonable design in 2025. It got typed output right — making a model return validated objects was not yet standard practice. Its ceiling is elsewhere: an agent cannot ask for anything, so what is not in the prompt does not exist to it.
- What it got rightoutput_type makes each agent return a validated Pydantic object, instead of free text somebody has to parse with regexes.
- The ceilingData is delivered once, up front. Wanting a specific sentence out of the 10-K, or a different discount-rate assumption, leads nowhere.
- The consequenceThe report is fixed at 13 chapters, because the pipeline is one-directional — it cannot decide what to do next from what it just saw.
04 · V2 architecture
Layered, with a tool boundary
V2 separates deciding what to do from doing it. The lead agent only orchestrates and judges; five role sub-agents carry the steps; and every number has to cross the tool boundary and be produced by a deterministic operator. A model cannot write a valuation out of thin air.
- 26 deterministic operatorsDCF, DDM, LBO, WACC, Monte Carlo, comparables, sum-of-the-parts, residual income — each a pure function with unit tests.
- 8 conversational toolsquery_financial_data · ask_filings · run_monte_carlo · run_backtest · find_reports · diff_reports · query_coverage_universe · activate_skill
- 7 pipeline toolsGenerated from the pipeline registry, so a whole analysis is something the agent calls: run_equity_research · run_dcf_valuation · run_lbo_analysis · run_ddm_valuation · run_comps_analysis · run_earnings_analysis · run_ic_memo
- ModelRetryWhen validation fails, the reason goes back to the model and it tries again — instead of failing the run or passing bad data downstream.
- 7 domain skill packsEquity research, financial analysis, investment banking, private equity, wealth management, plus two vendor packs — loaded on demand by activate_skill.
05 · The pipeline
The eight steps behind one report
This is the most concrete layer of “how we built it”. Each step is bound to one role agent, produces a validated structured object, and passes it forward as context — which is why a failure lands on a named step rather than on “report generation failed”. This is one of seven pipelines, and the lead agent reaches it the same way it reaches anything else: by calling run_equity_research.
| # | Step | Role | What it does |
|---|---|---|---|
| 1 | data_collection | data | Pull filings, prices and statements; cross-check the sources, emit structured financials |
| 2 | catalyst_analysis | data | Identify earnings dates, guidance changes and other event drivers |
| 3 | peer_analysis | analysis | Screen peers and normalise comparables (xbrl_aligned_comps) |
| 4 | financial_modeling | modeling | DCF / DDM / LBO / comparables — all computed by operators, then cross-reconciled |
| 5 | ownership_governance_analysis | analysis | 13F institutional holdings, Form 4 insider trades, governance structure |
| 6 | technical_analysis | modeling | Price momentum, volatility and technical signals |
| 7 | thesis | synthesis | Combine all of the above into an investment thesis and a rating |
| 8 | report | report | Write up: a 13-chapter report in which every number carries its origin |
The operators argue with each other. At step 4, a method whose result strays too far from the cross-method median is flagged as an outlier rather than quietly averaged in. From a real production log: ev_ebitda mid $261.20 deviates 42% from cross-method median $184.43 — flagged as outlier. That check lives in code; it does not depend on a model noticing.
06 · Why the design matters
Every number traces back to what produced it
In financial research, “where did this number come from” is the audit floor. A fair value written by an LLM cannot answer it — there are no intermediate steps, only tokens. In V2 every number has a retraceable operator chain, and the model's job is what sits outside that chain.
- AuditableEvery number traces to the operator, the inputs and the source version that produced it. In a compliance setting that is a requirement, not a bonus.
- ReproducibleSame inputs, same operators, same number. The model's temperature cannot reach it.
- Cross-checkedThe same field is fetched from both FMP and yfinance and flagged when they diverge past a threshold — production has logged a cash figure differing by 37%.
- InteractiveBecause the DCF is a function, a user can change an assumption and recompute. V1 cannot; its valuation section is prose that has already been written.
07 · The comparison
Four differences you can verify
“Better” is not worth claiming if it only means it feels better. Each row below corresponds to a structural difference.
| V1 · Report Studio | V2 · Research Desk | Why it is a difference | |
|---|---|---|---|
| Where numbers originate | a model doing arithmetic in tokens | 26 deterministic operators | Arithmetic errors and invented figures are structurally off the number path |
| Can it ask for more | No — nothing exists outside the prompt | 15 tools: query filings, run a Monte Carlo, diff two saved reports, dispatch a whole pipeline | The information that decides the next step can be obtained during the run, not staged in advance |
| How it is organised | 8 peer agents, none aware of the others | lead + 5 roles, an 8-step pipeline, each step bound to a role | Prompts can be scoped per role, and a failure localises to one step |
| Failure handling | validation fails → the step fails | ModelRetry returns the reason and the model retries | Bad data does not slip silently downstream — the most common failure in agent systems |
| What comes out | a fixed 13-chapter PDF | a traceable desk where assumptions can be changed and recomputed | Research is iterative; a one-shot document does not serve it |
Honestly, V2 is worse in places. It is slower and more expensive — layering plus tool calls means more model round-trips. Its dependency tree is heavier, and a cold start has to warm the compute engine and a symbol index. And for “just give me a standard report”, V1's one-directional pipeline is still faster, cheaper and more predictable. So the two coexist as two ways of working, not as an old version and its replacement — which is why this site keeps both entrances.