Source: Chen, Cho, Dou & Lev (2022) — “Predicting Future Earnings Changes Using Machine Learning and Detailed Financial Data”
At first this sounds like a technical finance question. Then you look closer and realise it is also a question about attention, incentives, and whether people are using information or just being impressed by it.
The standard approach to using accounting data in finance research is to construct a small set of summary statistics — return on assets, accruals, leverage ratios — and run a linear regression. This works reasonably well, but it involves substantial data reduction. By the time you’ve aggregated everything into five or six ratios, you’ve thrown away most of what was in the original filing. This paper asks what happens if you don’t throw it away. Using machine learning on high-dimensional, detailed financial data — granular line items from financial statements rather than aggregated ratios — the researchers predict the direction of one-year-ahead earnings changes. The models’ out-of-sample predictive performance is significantly better than conventional logistic regression approaches using small sets of accounting variables. The key finding: the area under the ROC curve ranges from 67.5% to 68.7%, compared to roughly 50% for a random guess. Hedge portfolios based on these predictions generate annual size-adjusted returns between 5% and 9.7%. What I find instructive here is the argument about data aggregation. Accounting standards require firms to report highly detailed information — individual balance sheet line items, revenue segments, specific expense categories.
In plain English, that is why the result matters beyond the chart. It changes where people should look, what they should question, and which comfortable assumption probably needs to be retired.
So I would not read this as a neat technology story. It is a messy information story. The tools improve, but people still have to decide what deserves attention. Annoying, but true.