New study shows how AI can expose inconsistencies in how logo trademarks are judged
New research by NYU Law’s Barton Beebe and Jeanne C. Fromer, co-authored with Vanderbilt Law School’s David Stein, delivers the first large-scale computational study of how the US Patent and Trademark Office (PTO) handles image trademarks—and their findings suggest that decades of logo registrations have been built on a surprisingly shaky foundation.
The article, “Trademark Law’s View of the Image: A Computational and Empirical Analysis,” uses AI to examine 2.9 million image-based trademark applications filed with the PTO between 1985 and 2023. Beebe, John M. Desmarais Professor of Intellectual Property Law, and Fromer, Vice Dean for University Partnerships and Walter J. Derenberg Professor of Intellectual Property Law, are co-directors of NYU Law’s Engelberg Center on Innovation Law & Policy.
The authors’ starting point is a long-standing puzzle at the heart of trademark doctrine: Judges and examiners have generally treated images as impossible to analyze in any rigorous way. As J. Thomas McCarthy—Professor Emeritus at the University of San Franscisco School of Law and a leading expert in the field of trademarks—expressed it, when it comes to visual similarity, “all one can say is ‘I know it when I see it.’” This subjective standard has persisted, despite the fact that image-only marks—logos, icons, and stylized designs without words—have come to make up a large and growing share of the trademark registry. In some industries, like clothing and footwear, image marks account for more than a third of all registrations.
To test whether the PTO’s case-by-case judgment calls appear to reflect consistent standards for review, the authors built machine-learning models that place millions of registered logos into a shared “similarity space,” allowing them to measure not just whether two marks resemble each other, but in what way—for instance, whether two logos both depict a dog, versus two logos that merely occupy the same stylistic territory.
Using that tool, the researchers discovered what they describe as systematic inconsistency. When the PTO refuses an application because of similarity to an already-registered mark, often it has already approved other marks that are objectively more similar to that same registered mark—and similar in the same way. In roughly half of PTO refusals, the authors found at least 100 previously registered marks in the same product category that were closer matches to the cited mark than the rejected applicant’s logo, yet faced no serious obstacles to registration.
One illustration involves a dog logo for the Northern Illinois University Huskies, which the PTO rejected as too close to the University of Washington’s Huskies mark. Running the numbers, the authors show that Northeastern University’s Huskies logo was more similar to Washington’s mark than Northern Illinois's logo was—and similar in the same visual and conceptual respects—yet Northeastern’s mark had cleared registration. The comparison, multiplied across the dataset, points to an examination process that does not apply a stable rule from one case to the next.
“What these AI tools have laid bare is that the review process as it currently stands is uneven at best,” says Fromer. “If current trends toward the proliferation of image marks persist, the PTO would do well to have more systematic and consistent ways of examining image marks. It might also consider responsible ways consistent with trademark’s goals of augmenting its current process with technological tools that could lead to more objectively equitable outcomes.”
The article also finds that the opposite problem has emerged on the question of inherent distinctiveness. PTO policy purportedly seeks to limit the protection of generic terms such as “furniture” for furniture in favor of less descriptive terms like “Ikea” for furniture. Over the decades studied, registered image marks have grown steadily more descriptive of the goods they represent, yet the PTO has rarely refused an image mark on distinctiveness grounds. That combination—being quick (or at least inconsistent) in raising similarity objections, while being slow to raise distinctiveness objections—is, the authors argue, exactly backwards from what a coherent trademark policy would produce.
Drawing on the finding, Beebe, Fromer, and Stein propose three doctrinal fixes: adopting an explicit “sight and meaning” test for image similarity that weighs both visual appearance and conceptual meaning, rather than privileging one over the other depending on the context; creating a hybrid standard for assessing when an image mark is inherently distinctive; and pushing courts and the PTO to investigate—rather than assume—how large the supply of alternative “image synonyms” really is for businesses seeking a distinctive look.
The authors are careful to frame computational tools as a diagnostic aid rather than a replacement for human judgment, cautioning against over-reliance on machine-learning models in high-stakes trademark disputes. “What our research shows,” Beebe says, “is that we now have the technology to conduct a more rigorous, systematic analysis that we can couple with qualitative human review to ensure that the overall judgments made are as fair and impartial as possible.”