The contested history of coffee cupping stretches back well before any written protocol, through Ottoman trading posts and San Francisco commodity floors, all the way to the Q Grader sitting across a cupping table in Addis Ababa. Most accounts start in the 1890s with a green coffee broker named Clarence E. Bickford and end with the SCA Cupping Protocol as the obvious conclusion. That framing is clean, linear, and incomplete.
What the standard narrative skips is messier and more interesting: the 50-year near-death of cupping as a practice, the unresolved tension between human perception and the myth of objectivity, and the question of who built the scoring systems that now determine what a farmer in Guatemala gets paid. The real story doesn’t end with standardization. It starts there.
Key Takeaways on the History of Coffee Cupping
- Coffee tasting existed informally for centuries before any written protocol; what changed in the 1890s was commercialization, not invention.
- Clarence E. Bickford and Hills Bros. proved that taste could justify price premiums, making cupping a commercial tool before it was a quality science.
- B.D. Balart’s 1932 blind-tasting procedure was the first serious attempt to remove individual bias from evaluation, predating the SCA by six decades.
- A 50-year near-silence between 1932 and Ted Lingle’s 1984 handbook shows how close cupping came to extinction under the weight of instant coffee and industrial roasting.
- The SCA’s 100-point scale and Q Grader certification turned cupping into a global market mechanism, not just a quality tool.
- Cupping scores are built on human perception that no protocol fully standardizes, and the standards themselves reflect the interests of the consuming-country actors who wrote them.
Before the Cup: How Coffee Was Judged Before Modern Cupping
Coffee tasting, in some form, existed long before anyone wrote a protocol for it. For centuries, the merchants moving green coffee beans through Yemen, the Ottoman Empire, and later through European trading houses evaluated their product with their eyes: screen size, color, density, and the absence of visible defects. These were the only metrics that traveled reliably across languages, borders, and time. No one had agreed on what a good cup was supposed to taste like, and in a commodity market, nobody needed to.
That visual-first system was practical, not lazy. When you’re moving sacks of coffee across the Red Sea on a ship with no refrigeration, what you care about is consistency and absence of damage. Size and color are proxies for those things. Taste is a different problem entirely – one that requires a shared vocabulary, a controlled environment, and the luxury of caring about nuance. The early coffee trade had none of that.
Still, the idea that no one tasted coffee before the 1890s strains credibility. Merchants in the Ottoman Empire trading coffee for four centuries were not buying blind. Informal tasting, the kind passed between traders in a back room or across a market stall, almost certainly existed as a practical necessity. It just wasn’t written down, standardized, or given a name. The limit wasn’t the absence of sensory judgment. The limit was the absence of a replicable, teachable method.
This is where the definition of coffee cupping gets philosophically complicated. If cupping means any act of evaluating coffee by taste, then it’s ancient. If it means a structured, bias-resistant procedure that produces transferable scores, then it’s modern. The standard history conflates the two, which is why it keeps arriving at the wrong starting line.
One source among the accounts surveyed offers a detail that none of the others bother to examine: that cup-testing was an unorthodox art passed down by word of mouth for more than 400 years before any Western formalization. The claim is unverified and the source doesn’t press it. But the implication is significant. If that oral tradition existed, then what Bickford and his contemporaries did in 1890s San Francisco wasn’t invent cupping. It was commercialize it. Standardize it. And in doing so, write its earlier practitioners out of the record entirely.
That distinction – between invention and institutionalization – is the lens through which this entire history should be read.

The San Francisco Breakthrough: Bickford, Hills Bros., and the Birth of Cup Quality
Clarence E. Bickford sits at the center of the standard origin story, and the credit is at least partially earned. As a green coffee broker working in San Francisco in the 1890s, Bickford made a move that was genuinely radical for its time: he started evaluating coffee by what it tasted like in the cup, not just what it looked like in the sack. In a trade built on visual commodity grading protocols, that was a provocation.
The logic behind it was commercial, not romantic. Visual inspection couldn’t tell you whether a smaller, high-grown bean would produce a better cup than a larger, lower-grown one. Hills Bros., the San Francisco roasting company that became an early adopter of cup testing, grasped this immediately. If you could demonstrate through taste that your coffee was superior, you could justify a price premium. Cupping wasn’t born as a quality movement. It was born as a pricing argument.
The next documented step forward came four decades later. In 1932, B.D. Balart published “Removing the Guess Work from Coffee Cupping” in the Tea & Coffee Trade Journal, and the title tells you everything about what problem he was trying to solve. Bickford and Hills Bros. had proven that taste mattered commercially. Balart’s contribution was trying to make the evaluation of taste something other than one man’s opinion. He introduced blind tasting as a procedural safeguard and catalogued 20 distinct cup defects, giving cuppers a shared vocabulary for what was wrong with a coffee rather than relying on vague impressions.
Between Bickford and Balart, William Ukers published All About Coffee in 1922, which stands as the first printed reference to the term “cup-test.” That’s a meaningful anchor: it’s the moment cupping entered the written record as a named practice, not just a habit of particular traders.
The causal chain here is clear enough. Visual grading couldn’t differentiate quality. Bickford and Hills Bros. demonstrated that taste could create commercial advantage. Balart attempted to make the procedure systematic and transferable so that advantage could be taught, scaled, and defended against subjectivity.
What the historical record is not clear about is whether Bickford deserves sole credit, whether Hills Bros. was the true catalyst, or whether Balart’s 1932 article is actually the first genuine formalization of the practice. The sources don’t agree. Some treat Bickford as the inventor. Others treat Hills Bros. as the institutional force that made it stick. Still others push the true birth of systematic cupping all the way to Ted Lingle’s 1984 handbook, which we’ll get to shortly. These aren’t minor footnotes. They reflect fundamentally different definitions of what cupping is: informal tasting, blind procedure, or teachable science. The history you accept depends on which definition you start with.
How Cupping Nearly Disappeared Between 1930 and 1980
Coffee cupping as a documented practice goes nearly silent after Balart’s 1932 article. For fifty years, the written record offers almost nothing. That silence isn’t accidental, and it isn’t just a gap in the archive.
The forces that filled those decades actively worked against sensory nuance. World War II brought coffee rationing to the United States, which trained an entire generation of consumers to accept whatever was available. After the war, the industry’s answer to mass demand was instant coffee and the consolidation of industrial roasting at scale. In that world, the relevant question wasn’t whether a Kenyan lot had more citric acidity than an Ethiopian one. It was whether you could fill enough cans to meet supermarket orders. Cupping, as a tool for differentiating individual lots, had no commercial function in that economy.
The specialty coffee movement of the 1970s and 1980s changed the equation. A small but growing number of roasters and importers started treating coffee as something other than a commodity, and that required a way to talk about quality differences between origins, farms, and processing methods. Erna Knutsen gave the movement its name in 1974 when she coined the term “specialty coffee” in the Tea & Coffee Trade Journal. Her company, Knutsen Coffees, used cupping as a selection tool for small, high-quality lots at a time when most of the industry had forgotten the practice existed. She didn’t publish a manual. She built a template through example, and a generation of buyers watched and learned.

The turning point came in 1984 when Ted R. Lingle published the Coffee Cupper’s Handbook through the Specialty Coffee Association of America. What Lingle did was deceptively simple and operationally enormous: he took what had been an oral, master-apprentice tradition and converted it into a printed, step-by-step methodology. For the first time, you didn’t need to apprentice under someone who knew how to cup. You could read the book.
That shift mattered commercially as much as technically. A specialty coffee market growing at any real speed cannot rely on knowledge that lives in the heads of a few experienced buyers. It needs a teachable skill. Lingle’s handbook was therefore as much a business infrastructure document as a sensory one. Without it, the specialty market that emerged in the 1990s and 2000s couldn’t have trained cuppers fast enough to function.
The fifty-year silence, properly understood, is not a gap in cupping’s history. It’s evidence of how close the practice came to extinction, and how contingent its revival was on a cultural shift that almost didn’t happen.
Standardizing the Senses: SCA Protocols, Flavor Wheels, and the Q Grader
The SCA Cupping Protocol, formalized in 1999, is the document that turned cupping from a specialty industry practice into a global standard. It fixed the variables that had always made cupping results hard to compare: brew ratio, water temperature, grind size, evaluation timing, and the structure of the scoring form. For the first time, a cupping result in Portland, Oregon carried the same procedural weight as one in São Paulo or Nairobi. Producers, buyers, and roasters finally had a shared language for quality.
The coffee tasting wheel evolved in parallel. The SCA’s first flavor wheel appeared in 1995, and it was a useful starting point: a circular map of broad flavor categories that gave cuppers a common vocabulary. But it was built on intuition and industry consensus, not sensory science. The 2016 revision changed that fundamentally. A collaboration between the SCA and the World Coffee Research organization brought together 72 sensory experts and produced a lexicon of 110 distinct flavor attributes, published in the Journal of Food Science. Each attribute was anchored to a physical reference standard, a specific chemical compound or food product that a cupper could actually smell or taste before evaluating a coffee. The gap between the 1995 wheel and the 2016 wheel is the gap between an industry vocabulary and a scientific instrument.
The Coffee Quality Institute (CQI), founded in 1996, took standardization one step further by creating a credentialing system. The Q Grader certification program required candidates to pass a battery of taste tests calibrated to SCA standards, using lab-grade equipment with 0.01-gram scale precision and a 10% grind tolerance. Pass the tests, and you became a certified Q Grader: one of the people authorized to produce official cupping scores that carry weight in commercial transactions.
As of recent counts, there are approximately 8,939 Arabica Q Graders and 543 Robusta Q Graders worldwide, totaling around 9,482 certified Q Graders globally. – Source: Gafei Coffee Knowledge
That number sounds large until you consider how many coffee transactions happen each year. Q Graders are the global credentialing body for cupping, and their scores feed directly into market mechanisms. The SCA’s 100-point scale defines “specialty” at 80 points and above. Cup of Excellence auction prices depend on Q Grader panel scores to set the lot rankings that determine what buyers will bid. The certification isn’t just a professional credential. It’s a gatekeeping function with real economic consequences downstream.
The Uncomfortable Truth: Subjectivity, Power, and What Standardization Obscured
Standardization solved a genuine problem. It also created a new one that the industry is only beginning to reckon with honestly.
The Persistent Subjectivity of Cupping
Sensory perception is not a fixed instrument. The founding promise of modern cupping, the one Balart articulated in 1932 and the SCA refined through the 1990s, was that bias could be systematically reduced through blind tasting, fixed protocols, and calibrated equipment. That promise was real. Blind tasting does reduce certain kinds of bias. Fixed protocols do improve reproducibility. Calibration sessions do bring individual cuppers closer together.
But they don’t eliminate the underlying variability in human perception. Aroma and flavour detection thresholds differ between individuals by orders of magnitude. Training reduces some of that variance, but it doesn’t erase it. Two Q Graders cupping the same coffee on the same day in the same room will not always produce the same score. The question of how much that variance matters, and when it matters most, is one the industry rarely discusses publicly because the answer is inconvenient for a system built on the authority of certified scores.
The early methods Balart used, a defect list and a structured but simple tasting procedure, had a kind of honest humility to them. They acknowledged limits. Modern protocols, by layering on precision equipment and certification requirements, have created an impression of scientific objectivity that the underlying biology of human tasting doesn’t fully support.
Chemical analysis and digital sensors exist as alternatives. Neither has replaced the human cupper, and not primarily for technical reasons. The market requires the social act of cupping to function: the shared table, the agreed vocabulary, the ceremony of evaluation. That social function is real and valuable. But conflating it with objective measurement is a story the industry tells itself, not a fact about the practice.
Economic Power and Performance: How Cupping Scores Shape Producer Livelihoods
Cupping scores don’t just describe coffee. They set prices. And the people who built the scoring systems, the vocabulary, and the credentialing infrastructure are overwhelmingly located in coffee-consuming countries, not coffee-producing ones.
Hills Bros.’ use of cup testing to justify premium prices for high-grown beans is remembered as an innovation in quality assessment. It was also the construction of a new mechanism for determining value in a trade relationship where the power was already asymmetric. The farmer growing the beans had no seat at the table where the scoring criteria were written.
A peer-reviewed study in the Journal of Agricultural and Applied Economics analyzing Cup of Excellence auction data found that while sensory quality scores do contribute to specialty coffee prices, symbolic attributes such as geographic origin and buyer market conditions are significant price drivers. The results suggest that value capture is heavily influenced by downstream consuming-country actors, pointing to a structural asymmetry in how cupping-based premiums are distributed across the supply chain.
The finding matters because it names something the celebratory history of cupping never does: the scoring system is not a neutral instrument applied to coffee from outside the market. It is itself a market mechanism, built by specific actors with specific interests, and its outputs determine how much a smallholder farmer in Ethiopia or Colombia gets paid for a year’s work.
This doesn’t make cupping corrupt. It makes it human. Systems built by people with interests reflect those interests. The problem isn’t that the SCA protocol exists. The problem is presenting it as a neutral technical standard when it is, in part, a political settlement about whose palate counts and whose definition of quality governs the price.
“Performative cupping” captures the far end of this dynamic: the elaborate tableside ceremony staged for buyers or marketing purposes, where the evaluation ritual substitutes for actual evaluation. When cupping becomes spectacle, it drifts furthest from Bickford’s original provocation and Balart’s original intent. It becomes theater in the vocabulary of quality rather than the practice of it.
The Living History: Why Cupping’s Origins Matter to Every Coffee Drinker Today
Coffee cupping is not a historical artifact. Every number on a specialty coffee bag, every Cup of Excellence lot description, every flavor note on a third-wave roaster’s website is a direct output of the system this history built.
When a coffee scores 90+ at a Cup of Excellence auction, that number is the end of a chain that runs from Bickford’s cup test in 1890s San Francisco through Balart’s defect list, through Lingle’s handbook, through the SCA protocol, through the Q Grader certification, and through the 2016 flavor wheel. Every link in that chain was built by specific people with specific problems they were trying to solve. Understanding the chain tells you what the number actually measures and what it doesn’t.
When a barista reaches for a flavor descriptor, calling a coffee “jasmine-forward” or “red currant,” they’re not just describing a sensory experience. They’re participating in a vocabulary that was assembled over 150 years of negotiation between traders, scientists, roasters, and certifying bodies. That vocabulary is powerful and genuinely useful. It’s also incomplete, and the incompleteness reflects the interests of the people who built it.
The line from Hills Bros.’ premium pricing to today’s direct trade and relationship coffee models is shorter than it looks. The commercial logic hasn’t changed: cupping scores justify price differentials. What has changed is that producers in Ethiopia, Colombia, Kenya, and elsewhere are increasingly building their own cupping labs, training their own Q Graders, and entering the evaluation conversation as participants rather than subjects. That shift is the most consequential development in cupping since the SCA protocol. It’s happening now, and the SCAA cupping protocol that governs so much of this evaluation is being interrogated by the very producers it was designed to assess.
The growing calibration movement within the coffee industry, the push for sensory science research that acknowledges inter-cupper variability, and the producer-led quality assessment programs emerging in origin countries are all responses to the same underlying tension this history exposes. The question of whose palate counts is not settled. It’s being renegotiated in real time.
The history of cupping teaches one thing clearly: standards are made by people with interests. Every time you read a cupping score, a flavor note, or a quality claim on a coffee bag, you’re reading a document produced by that history. Knowing the history doesn’t make the coffee taste different. It makes you a sharper reader of what you’re being told about it.
Frequently Asked Questions About the History of Coffee Cupping
Did coffee producers ever have input into how cupping standards were developed?
Historically, no. The SCA protocol and Q Grader certification were built almost entirely by consuming-country actors – brokers, roasters, and buyers in the US and Europe. Producer-country participation in setting evaluation standards is a recent and still-incomplete development.
Why didn’t Balart’s 1932 system become the global standard immediately?
It didn’t scale because it stayed in a trade journal. A published article can describe a method, but it can’t train people. Without an institutional body to enforce calibration or a handbook to teach from, the method couldn’t spread beyond the readers who happened to find it.
How does the SCA’s 80-point specialty threshold actually affect a farmer’s income?
A coffee that scores 79 points sells as a commodity at futures-market prices. At 80 points, it qualifies as “specialty” and can command a significant premium. That one-point boundary, set by a protocol written in consuming countries, has direct consequences for smallholder income in producing countries.
What’s the difference between a Q Grader score and a roaster’s own cupping notes?
A Q Grader score is produced using CQI-standardized procedures and equipment, calibrated to SCA criteria, and carries official commercial weight. A roaster’s cupping notes are an internal evaluation with no required standardization. Both are useful, but only one can officially define a coffee as “specialty.”
How reliable are cupping scores between different trained tasters?
More reliable than untrained tasting, but less reliable than most buyers assume. Studies on inter-cupper reliability show meaningful score variance even among certified Q Graders cupping the same coffee. Calibration sessions reduce this, but they don’t eliminate the underlying variability in human perception.
Why did Erna Knutsen matter to cupping’s revival if she didn’t publish a protocol?
Knutsen proved there was a market for quality differentiation at a time when the industry had nearly abandoned the concept. She built a buying practice around cupping before a handbook existed to teach it. That proof of commercial viability created the demand that Lingle’s 1984 handbook then met with a teachable system.
What is “performative cupping” and why does it matter?
Performative cupping is the use of the cupping ritual as marketing theater rather than genuine evaluation. It matters because it hollows out the practice’s original purpose: when the ceremony substitutes for the judgment, scores lose their grounding in actual sensory analysis and become a brand signal instead.
Could digital sensors or chemical analysis ever replace human cuppers?
Technically, chemical analysis can identify flavor compounds with more precision than any human nose. But the market hasn’t adopted it because cupping serves a social function: a shared evaluation ritual that creates trust between buyers and sellers. Replacing the human taster would require replacing that social infrastructure, not just the sensory tool.
References
- What Explains Specialty Coffee Quality Scores and Prices: A Case Study from the Cup of Excellence Program | cambridge.org
- Coffee Q Grader Statistics | en.gafei.com





