Investor's Last Exam by ProphetLab Main navigation主导航
Language语言
Colour stamp配色

Investor's Last Exam · financebench-v9-60 · 57 tasks

Equity research tasks and the financial environment 股票研究题目与金融环境

Twenty task types across five equity research categories, instantiated as 57 tasks with access to financial data up to each task’s own cutoff. 五大类股票研究问题、二十种题型,共 57 道题;每道题只使用其数据截止时间前可获取的金融信息。

Investor's Last Exam is ProphetLab’s dataset and research environment for financial forecasting and equity research. Tasks connect evidence, financial reasoning, scenario analysis and probabilistic forecasts. Investor's Last Exam 是 ProphetLab 面向金融预测与投资研究构建的数据集与研究环境,考察从证据检索、财务推理到情景分析与概率预测的综合能力。

01 / the task set题目集合

Five categories, twenty task types五大类问题,二十种题型

The five categories cover opportunity discovery, evidence gathering, fundamental forecasting, investment analysis and ongoing monitoring. Each question is evaluated independently and has its own answer requirements and scoring criteria. 五类问题覆盖机会发现、信息获取、基本面预测、投资分析与持续跟踪。各题独立作答和评分,并分别设置答案要求与评分标准。

Expand a task type to read a sample question. Depth labels summarise the median number of evidence links needed for a sufficient answer; they are not difficulty ratings. The five depth groups contain 19, 19, 8, 10 and 1 tasks respectively. 展开题型可查看代表性题目。深度标签表示形成充分答案所需证据关联步数的中位数,不等同于难度评级。五档深度分别包含 19、19、8、10、1 道题,最高档目前仅有一个样本。

Research-informed task design面向投研需求的题目设计

ProphetLab’s AI post-training and investment research teams develop the tasks with input from investment practitioners on questions and scoring criteria. Each task has its own answer requirements, covering conclusions, calculations and supporting evidence. 题目由 ProphetLab 的 AI 后训练与金融投研团队建设,一线投资从业者参与问题与评分标准设计。每道题分别设置答案要求,考察结论、计算过程与支持证据。

Practitioners help select tasks that analysts can complete but remain challenging for frontier models, including portfolio adjustments under risk constraints, implied-expectations analysis and pre-earnings forecasts. 从业者参与筛选分析师能够完成、但前沿模型仍具挑战的任务,例如风险约束下的持仓调整、股价隐含预期计算与财报发布前的业绩预测。

Research across companies and periods跨公司、跨期间的长程研究

Tasks require multiple rounds of evidence gathering and validation across companies and reporting periods. Agents connect financial disclosures with supplier and customer relationships, supply and demand, and competitive conditions to revise their forecasts. 任务要求模型跨公司、跨财务期间开展多步取证与反复验证,将财务披露与上下游关系、供需变化和竞争格局联系起来,形成并修正预测。

For example, an external advertising-market signal may inform an adjustment to a company’s revenue guidance, with the effect supported by explicit calculations. 例如,以外部广告市场信号调整公司的收入指引,并用明确的计算说明这一信号如何影响预测。

The 171 completed runs average 22.7 main-agent responses per task, excluding reader subcalls and API retries. 本轮 171 次完整作答,平均每题 22.7 轮主 Agent 回复,不含 reader 子调用与 API 重试。

02 / the data lake数据湖

Over one million financial documents百万级真实金融文档

The data lake brings together 3–5 years of financial information across A-share, Hong Kong and US markets: more than one million documents and 35 billion characters. Company filings, earnings transcripts, investor materials, financial statements and news are updated incrementally, with publication times recorded at the precision available from each source. 数据湖汇集过去 3–5 年 A 股、港股和美股的金融资料,文档超过 100 万份、字符逾 350 亿。公司公告、业绩会记录、投资者关系材料、财务报表与新闻等持续增量更新,并按数据源可提供的精度记录发布时间。

The data lake covers three markets and eight document types. This evaluation uses annual reports, earnings transcripts and news for ten US-listed companies. 数据湖覆盖三个市场、八类文档。本次评测使用十家美股上市公司的年报、业绩会记录与新闻。

03 / the freeze冻结

Point-in-time financial information按时点限定的金融信息

All 57 tasks can access annual reports, earnings transcripts and news within their permitted time range. The minimum evidence requirements vary with each question. 57 道题均可访问各自截止时点前的年报、业绩会记录与新闻。每道题所需的最低证据数量,取决于具体研究问题。

Forecasts and subsequent disclosures预测与后续披露

Forecast questions require conclusions based on information available at the task cutoff. Subsequent disclosures provide realised outcomes for comparison. 预测题要求模型依据截止时点前的信息形成判断,后续披露提供可对照的实际结果。

Access controlled by the environment由环境控制数据访问

Each task places the agent at a specific historical moment. Publication times and visibility boundaries determine which documents it can retrieve, excluding later disclosures from the information available for research. 每道题将 Agent 置于一个具体的历史时点。环境依据发布时间与可见时间边界限定可检索文档,研究过程中无法通过工具获取该时点之后的披露。

The same task-level information boundary applies to search, reading and file downloads for all three models. 同一题目的检索、阅读和文件下载遵循一致的信息边界,三个模型采用相同访问规则。

Recorded evidence and retrieval steps support reproducible historical research. The same point-in-time data environment can also support backtesting and model post-training. 记录下来的证据与检索过程便于复核历史研究;同一套按时点限定的数据环境,也可用于历史回测与模型后训练。

04 / Research environment研究环境

A shared set of research tools统一的研究工具

Each agent uses the same tools to find evidence, analyse financial information and submit an answer. 每个 Agent 使用相同的工具查找证据、分析财务信息并提交答案。

Evidence is limited to each task’s information cutoff. Calculations run in an isolated environment without public internet access. 可用资料受每道题的信息截止时间限制,计算在无公网访问的独立环境中完成。

05 / the tool use工具使用

Tool use across models模型的工具使用情况

Compare how often each model searches, reads, downloads, computes and submits across the 57 tasks. 比较各模型在 57 道题中检索、阅读、下载、计算与提交的工具使用情况。

Normalisation归一化方式

3 models · field mean on top3 个模型 · 首行为全场均值

06 / Capability differences能力差异

Where model capabilities differ模型能力差异体现在哪里

Selected equity research cases reveal differences in financial accuracy, forecasting judgment and the ability to recover evidence when a research path stalls. 通过精选投研案例,比较模型在财务公式、收入预测逻辑,以及取证受阻后切换工具方面的能力差异。

Financial formula accuracy财务公式的准确性

In the Coca-Cola reverse-valuation task, Ling subtracts non-controlling interests when calculating enterprise value, while GLM and DeepSeek add them. This sign difference also changes the implied growth and margin estimates. 可口可乐反向估值题中,Ling 在计算企业价值时减去了应加回的少数股东权益,GLM 和 DeepSeek 则正确加回。一个符号的差异,也会改变后续隐含增速与利润率的估计。

The task uses enterprise value = market capitalisation + debt − cash + non-controlling interests. Subtracting rather than adding the $2.106 billion interest understates enterprise value by $4.212 billion, lowering the cash-flow target in the reverse-valuation model. 这道题采用“企业价值=市值+债务-现金+少数股东权益”。将 21.06 亿美元少数股东权益由加项写成减项,会使企业价值低估 42.12 亿美元,进而降低反向估值所需匹配的现金流目标。

Reasoning behind a revenue forecast收入预测逻辑的合理性

In the Microsoft pre-earnings sample, Ling extrapolates from recent results using seasonality and growth assumptions, without anchoring revenue to management’s guidance. DeepSeek largely uses the guidance midpoint. GLM starts from that midpoint and makes an evidence-based upside adjustment. 微软财报前预测样例中,Ling 主要依据近期业绩、季节性与增长率假设外推,未以管理层收入指引为锚;DeepSeek 基本采用指引中值;GLM 则在指引基础上,结合证据作出上调。

GLM considers prior guidance beats, Azure capacity coming online earlier and Copilot growth, while also accounting for pressure on Windows and gaming. This is closer to an analyst’s evidence-calibrated approach: use guidance as a starting point, then assess whether current business conditions justify a deviation. A higher forecast alone does not make the reasoning stronger. GLM 同时考虑历史超指引表现、Azure 产能提前交付、Copilot 增长,以及 Windows 和游戏业务的压力。这更接近人类投研专家的分析路径:先以指引为起点,再结合当期经营情况判断是否应偏离指引,而不是仅作机械外推或直接取中值;关键是调整有无依据,而非预测是否更高。

Switching tools to recover evidence取证受阻后的工具切换能力

In the Meta forecast-update sample, the models need Alphabet’s advertising revenue as an external signal. When document reading fails to provide the required figures, GLM and DeepSeek switch to downloading the original filings and extracting the data locally, reconstructing aggregate advertising growth from Search, YouTube and Network revenue. Meta 预测更新样例需要用 Alphabet 广告收入作为外部信号。阅读工具未返回所需数字时,GLM 和 DeepSeek 转向下载原始披露并在本地提取数据,从搜索、YouTube 与广告网络收入重建整体广告增速。

Ling continues searching and reading but does not use the download route in this sample, ultimately relying on an approximate growth figure. GLM and DeepSeek more effectively turn a blocked reading path into verified source data; Ling’s response exposes a weakness in adapting its tool strategy. Ling 则继续搜索与阅读,在这条样例中未使用下载路径,最终采用近似增速。相比之下,GLM 和 DeepSeek 能更有效地通过工具切换找回并核对原始数据;Ling 在取证受阻后的策略调整上表现不足。

07 / one run, end to end一次完整运行

An agent at work: a recorded research workflowAgent 的真实工作流——以一条轨迹为例

Follow GLM-5.3 as it updates a Tesla forecast, from finding and reviewing sources to calculations and final submission. 以 GLM-5.3 更新特斯拉预测的一次作答为例,展示从资料检索、阅读取证到计算与提交的完整过程。

Column height柱高口径

Cumulative elapsed time appears below each round每轮下方标注累计耗时

The round-by-round record connects model responses with tool operations. Scoring assesses the final submitted answer; model API costs depend on token usage. 逐轮记录展示模型响应及对应的工具操作。评分依据最终提交答案,模型 API 费用则按 token 用量计算。

Model and tool time模型调用与工具执行耗时

The 163.6-second research loop includes 127.5 seconds of model calls, 25.0 seconds of tool execution and 11.1 seconds of orchestration overhead. Reading accounts for 16.8 seconds of tool time. Model calls dominate the total duration. 163.6 秒的研究过程中,模型调用占 127.5 秒,工具执行占 25.0 秒,调度等开销占 11.1 秒。其中,阅读工具耗时 16.8 秒。总耗时主要来自模型调用。

Corrections before final submission最终提交前的修正过程

In round 9, the first submit_answer attempt is rejected because the evidence array does not meet the required format. In round 10, the model corrects the fields and resubmits while preserving its research. The two rounds produce 2,422 and 2,017 output tokens respectively. 第 9 轮首次调用 submit_answer 时,evidence 数组不符合规定格式,提交被拒绝。第 10 轮模型保留研究内容,修正字段后重新提交。两轮分别产生 2,422 和 2,017 个输出 token。