← Back to feed
2026-06-24agentsreasoningdata

InvestPhilBench: A Multi-Layer Dynamic Benchmark for Evaluating Large Language Model Procedural Reasoning in Expert Investment Philosophy

Mingguang Chen, Bo Qu

PDF preview for InvestPhilBench: A Multi-Layer Dynamic Benchmark for Evaluating Large Language Model Procedural Reasoning in Expert Investment Philosophy
Read on arXiv →

Key claim

Models excel in fluency but struggle with procedural reasoning.

In plain English

Imagine you're trying to build a system that helps investors make decisions based on complex frameworks. Currently, many systems use large language models to assist with this, but they often fail to accurately replicate the nuanced decision-making processes that expert investors use. This can lead to poor investment choices because the models might sound convincing but lack the depth of understanding needed for real-world applications. This is what's called procedural deficit — the models can generate fluent text but don't always follow the right reasoning steps to arrive at sound conclusions.

To address this, the authors created InvestPhilBench, a benchmark that evaluates how well these models can reconstruct and apply expert investment frameworks. They developed a scoring system that includes various metrics to assess model performance, focusing on both fluency and procedural accuracy. The benchmark includes a wide range of investment principles and decision frameworks, allowing for a thorough evaluation of the models' capabilities.

The results show that while the models achieve high scores in generating text, they still fall short in accurately following the procedural steps that expert investors would take. This means that for anyone building investment tools using language models, it's crucial to recognize that high fluency doesn't guarantee sound decision-making. The findings suggest that future work should focus on improving the procedural reasoning of these models to better support investment decisions.

Novelty
8.0/10

Introduces a comprehensive benchmark for evaluating language models in investment contexts.

Reliability
7.5/10

Provides solid experimental validation with multiple metrics and a clear scoring pipeline.

Deep reliability assessment

The methodology supports the paper’s main benchmark-design claim: composite free-text scoring can look saturated while gate-level scoring still finds procedural failures. The model-performance claims should be treated as preliminary rather than leaderboard-grade because the sanity wave is closed-book, mixed-judge/confounded, and the de-confounded retrieval/oracle evaluation is deferred to v1.0.

Reproducibility

Partial: the paper describes a v0.6 release with 118 principle cards, 25 framework cards, 243 QA questions, BASP/FMDP/GRA metrics, and a 100-item expert-annotated gold set, but no public repository or project URL is visible in the provided text. The main multi-model leaderboard and full three-condition evaluation are explicitly future v1.0 deliverables.

Key figure

Figure 1 plots BASP composite scores by cognitive layer for four models, showing frontier models staying high across L1–L8 while cheaper Gemini models sit lower and collapse at L8, with no visible L3-to-L4 composite cliff.

Benchmark results

~InvestPhilBench v0.6 development splitBASP composite: 0.906vs lower provider-tier models, lowest reported BASP 0.438+0.468 absolute vs lowest reported model
~InvestPhilBench v0.6 development splitBASP composite: 0.932vs not specifiednot specified
~InvestPhilBench v0.6 questions with gold reasoning programsGate Reconstruction Accuracy: 0.77vs not specifiednot specified
~100-item expert-annotated gold setPearson correlation between BASP composite and human reference: 0.72vs not specifiednot specified