A Camera Can Now Grade How Fermented Your Tieguanyin Is
A 2026 study taught machine-learning models to read Tieguanyin's fermentation stage from ordinary photos and hyperspectral scans, graded against a two-of-three vote among tea masters, not a chemical threshold and not one expert's word alone.
A camera and a hyperspectral scanner, used together, can name which of six fermentation stages a batch of Anxi Tieguanyin is in, correctly, 94.44 percent of the time. That figure comes from a study published in Foods on 12 January 20261 by Yuyan Huang, Yongkuai Chen, Chuanhui Li, Tao Wang, Chengxu Zheng and Jian Zhao of the Institute of Digital Agriculture, Fujian Academy of Agricultural Sciences. It is the best result in the paper. The Order does not repeat a best result without its worst: other configurations the same team tested finished at 68.06 percent, which is wrong roughly one time in three.
Our own guide to oxidation sets out what the reaction is, how it runs, and what light and heavy oxidation do to a cup. It does not answer the question a maker faces with the leaf spread out in front of him. Is it done yet. That question, the paper says in its own introduction, "currently relies primarily on the experience and subjective judgment of tea masters," an approach it calls "highly subjective and uncertain, making it difficult to meet the demands of modern, standardized, and large-scale production." The researchers' own answer to that subjectivity was not to trust one master's word for it. It took three, watching the same leaf, with a stage counted as reached only once at least two of the three agreed. Once the leaf goes to heat for kill-green, the argument, whatever was left of it, is settled by force. Whatever oxidation had reached at that instant is what the tea is, permanently.
Six stages, timed to the hand and not the clock
The hardware is not the interesting part. The labels are.
The researchers did not divide fermentation into even percentages. They divided it into the maker's own actions. Stage one is before the first aeration. Stage two is the end of the first aeration. Stage three is the end of the second. Stage four is three hours into the third. Stage five is optimal fermentation, the point at which the leaf was judged ready. Stage six is over-fermentation, defined as one hour past that point.
Stage five and stage six are the same leaf, sixty minutes apart, and "judged ready" is the two-of-three vote described above, run again for each batch. That vote is the ground truth the model learns from. Not a chemical threshold, and not one person's palate either, but a committee's majority. So "94.44 percent accuracy" has a narrower meaning than it first sounds. The model was never checked against an objective measurement in the leaf, because for this question there is not yet one to check against. It was checked against whatever two masters out of three could agree on, and it matched that agreement 94.44 percent of the time at best.
The material was real production tea, Anxi Tieguanyin from Binghuai Family Farm in Longjuan Township, Anxi County, Fujian, sampled across three separate batches on 3 and 4 May 2024, during the actual fermentation runs rather than in a laboratory imitation of them.
What each sensor sees
Two instruments were pointed at the leaf.
The first is an ordinary RGB camera, the same three channels your eye has. Four hundred images per stage, averaged down into 40 datasets per stage. This is the mechanized version of looking: the reddening rim, the mottling of the leaf face, the retreat of green from the edge inward.
The second is a hyperspectral camera reading from 400 to 1000 nanometers across 224 bands. Where your eye collapses that entire range into three numbers, the instrument keeps 224 of them, including the near-infrared beyond 700 nanometers, which no eye registers at all. Forty hyperspectral samples per stage. Together, 240 datasets, split 168 for training and 72 for testing in a 7:3 stratified split.
Three model families were tried against the data: a Support Vector Machine, a Long Short-Term Memory network, and a Gated Recurrent Unit. Each was run three ways, on images alone, on spectra alone, and on fused combinations of the two.
The pattern in the results is more useful than any single score. Image-only models clustered tightly, 76.39 to 80.56 percent. Colour is genuinely informative and it plateaus fast. Hyperspectral-only models scattered wildly, 68.06 to 87.50 percent, the same sensor producing both one of the weakest results in the study and one strong enough to beat most fusion configurations, depending only on how the features were chosen. Raw spectral data carries more signal than colour does and drowns it in noise unless the reduction is done well. Fusion models ran 77.78 to 94.44 percent, and took the top spot: the best hyperspectral-only score still trailed three separate fusion configurations, 93.06, 93.06, and 91.67 percent.
The winner was a pipeline the authors label Pearson+Nor-SPA+L1+SVM: Pearson correlation and a normalized successive projections algorithm to select the informative wavelengths, L1 regularization to keep the model from leaning on all of them, and a Support Vector Machine to make the final call. The bands it leaned on hardest were 967, 942, 814, 784, 781, 503, 413 and 416 nanometers. Five of those eight sit past the red limit of human vision.
Neither sensor alone won. A taster does not judge with one sense either, and the fused model beating both of its own inputs is the closest thing in this paper to a compliment paid to the panel of masters it is meant to assist.
The part that is not proven
The Order's position on 94.44 percent is that it is a real number and a small one.
It means about one batch in eighteen is placed in the wrong stage, and that is the best pipeline of the many tested. If the misplacement falls between stage five and stage six, the consequence is a tea taken an hour late, the difference between a Tieguanyin that holds its florals and one that has gone soft and dull, not a rounding error in the cup. Choose your feature selection badly and you are at 68.06 percent, which is worse than useless, because a confident wrong answer is more dangerous than no answer.
The narrower limit is the sample. One farm. Two consecutive days. One spring. Three batches of one cultivar in one township. The paper does not claim the model transfers to another producer's aeration schedule, to autumn leaf, to a different cultivar, or to a yancha roast regime where the whole processing logic differs. Nothing here has been shown to survive a change of farm. Move the camera to a different hillside and there is, so far, no evidence the model still knows what it is looking at.
The authors frame the work as a step toward standardized, scalable quality control, not as a replacement for the craft. That framing is accurate and the Order will not inflate it. What the study demonstrates is that the information three tea masters read off the leaf, together, is at least partly present in measurable form, that some of it lives in wavelengths no human has ever seen, and that a machine can recover enough of it to be right most of the time on the tea it was trained on.
Of the 72 samples held back for testing, the best model placed 68 correctly.