Investor's Last Exam · financebench-v9-60 · 2026-09-02
Model rankings and capability differences 模型排名与能力差异
Compare overall performance, then examine differences across task categories, evidence depth and scoring dimensions. 先比较综合得分,再从题目类别、证据深度和评分维度了解各模型的优势与短板。
57 tasks · 1,106 scored criteria · one attempt per model · results as of 2 September 2026. 57 道题 · 1,106 条计分评分点 · 每个模型作答一遍 · 结果截至 2026 年 9 月 2 日。
01 / the score总得分
Overall model scores模型综合得分
The overall score is the average of weighted task scores across the selected market. This evaluation covers 57 tasks from 10 US-listed companies; A-share and Hong Kong evaluations are planned. 综合得分是所选市场各题加权得分的平均值。本次评测覆盖 10 家美股上市公司的 57 道题;A 股和港股评测正在规划中。
3 models · no human panel in this run3 个模型 · 本轮无人类对照组
02 / the full result set完整结果
Detailed model results模型详细结果
Compare overall scores, scores from each judge, task-level performance and resource usage. Click a column heading to sort. 对比综合得分、两名评审的得分、各题表现与资源消耗。点击表头可排序。
03 / by task category按题目类别
Five categories of equity research五大类股票研究问题
Scores across five equity research task categories show where each model is strongest and where it needs to improve. 五类投研题目的分项得分,展示各模型擅长的任务以及仍需提升的能力。
04 / category scores分类得分
Rankings by task category各类题目的得分排名
Compare model scores across all five categories, or examine a single category alongside each model’s overall score and rank. 比较三个模型在五类题目中的得分差异,也可查看单类题目的成绩,以及它与模型综合得分、排名的关系。
05 / by evidence depth按证据深度拆
Performance across evidence depths不同证据深度下的表现
Depth here is the blueprint's shortest sufficient path through the evidence graph — the fewest hops that can answer the question at all. A 1–3 hop task is one extraction with a citation; an 8–11 hop task chains filings, transcripts and news before the first number can be written down. 这里的「深度」指题目设计中足以作答的最短证据路径——最少要走几跳才能答出来。1–3 跳的题目是一次带出处的抽取;8–11 跳的题目要把年报、电话会记录与新闻串起来,才能写下第一个数字。
06 / every cell全矩阵
Model × task category模型 × 题目类别
Compare models within each task category, or compare a model across categories to identify its strengths and weaknesses. 按列比较同类题目中各模型的表现,按行比较同一模型在不同题目类别中的优势与短板。
07 / rubric categories评分类别
What the score is made of得分由什么构成
Task-specific criteria are grouped into seven dimensions for analysis. These dimension scores complement the overall score by showing the quality of conclusions, reasoning and evidence. 各题的评分点归为七个维度进行分析,分别观察结论、推理与证据等方面的质量,补充综合得分无法体现的能力差异。