By Seth Earley, Founder & CEO, Earley Information Science
────────────────────────────────────
Published: December 2025 | Article 7 of 10 in the *Scaling GenAI* Series
Last Updated: August 2026 | Version 1.1
────────────────────────────────────
Who This Is For: Taxonomists, information architects, and content strategists responsible for designing classification systems that support AI retrieval. Essential reading for anyone whose carefully designed taxonomy isn't delivering expected search or AI results.
Prerequisites: Familiarity with taxonomy and classification concepts; experience with enterprise search or content organization helpful.
────────────────────────────────────
Your taxonomy is probably a tree.
Products at the top. Hardware and Software as branches. Laptops and Desktops under Hardware. Operating Systems and Applications under Software. Clean. Logical. Hierarchical.
Now answer this question: where does a laptop with pre-installed software go?
It is hardware, but it is also software. It is a laptop that runs an operating system and comes with applications. In your carefully designed tree, this everyday product does not fit neatly into any category.
This is not a theoretical problem. This is why your GenAI retrieves the wrong content.
Rigid hierarchies force content into single categories when reality is multi-dimensional. When a user asks, "What laptops come with productivity software pre-installed?" a hierarchical taxonomy cannot connect hardware products to software bundles because they exist in different branches.
The outcome: AI provides incomplete answers, overlooks relevant content, or returns no results.
The solution is not a better hierarchy. It is abandoning hierarchy as the primary organizing principle.
The Hierarchy Trap
Hierarchical taxonomies made sense in the world of filing cabinets and library card catalogs. Physical objects can only exist in one place. A book sits on one shelf. A document goes in one folder.
But digital content does not have that constraint. A product spec can be simultaneously about laptops as a product category, about Windows as a platform, for IT procurement as an audience, relevant to a Q4 refresh cycle for timing, and related to security compliance as a use case.
Forcing that document into a single branch of a tree throws away most of its retrievable dimensions.
Hierarchies create three specific problems for GenAI.
Problem 1: Single-Path Retrieval
In a hierarchy, there is one path to each piece of content. If users do not navigate that exact path, they do not find the content.
For example, a user asks: "What are our remote work policies for the Austin office?" The content is filed under HR, then Policies, then Work Arrangements, then Remote Work. The challenge is that the query mentions "Austin office," a geographic dimension that does not exist in the hierarchy. The requested content exists, but the taxonomy makes it invisible.
Problem 2: Artificial Boundaries
Hierarchies create walls between related content that happens to sit in different branches.
Sales collateral about a product lives under Marketing. Technical specs for the same product live under Engineering. Support documentation lives under Customer Service. A customer asks a simple question: "What does Product X do and how do I set it up?" The AI would need content from three different branches, but the hierarchy treats them as unrelated. The retrieval is fragmented or incomplete.
Problem 3: Maintenance Nightmare
Hierarchies are rigid. When the business changes, new product lines, reorganizations, merged categories, the entire taxonomy must be restructured. Documents have to be reclassified. Links break. Search results degrade.
Every organizational change becomes a taxonomy change, and taxonomy changes ripple across thousands of documents.
The Adaptive Alternative: Faceted Classification
Faceted classification solves these problems by describing content across multiple independent dimensions instead of forcing it into a single hierarchy.
Now that laptop with pre-installed software can carry a product type of Bundle, categories of both Laptops and Application, a platform of Windows, an audience of Enterprise, a use case of Productivity, and relationships to Windows 11, the Office Suite, a Security Suite, and a Setup Guide.
The same product is discoverable through multiple paths. Ask about laptops, it appears. Ask about productivity software, it appears. Ask about enterprise bundles, it appears. Ask about Windows devices, it appears.
Faceted classification means content can be found however users think to ask for it.
Is-ness and About-ness: The Foundation of Content Models
Before you can build a faceted classification, you need to understand what you are classifying. This requires answering two fundamental questions.
Is-ness: What Is This Thing?
Is-ness defines the fundamental nature of a piece of content. If you handed it to someone, what would you call it? A policy document, a product specification, a troubleshooting guide, an FAQ, a procedure, a training module, a case study.
Is-ness is typically a controlled list. You do not have an infinite number of content types; you have a manageable set that your organization produces.
About-ness: What Makes It Different?
About-ness answers: if you had 1,000 documents of this type, how would you tell them apart? What piles would you sort them into?
For 1,000 policy documents, that might mean policy domain, audience, geography, effective date, compliance requirement, and related policies. For 1,000 troubleshooting guides, that might mean product line, product model, error code or symptom, severity level, required expertise level, and related procedures.
About-ness is where faceted classification lives. Each dimension of about-ness becomes a facet that makes content discoverable.
The Design Process
Start by identifying your content types, the is-ness. What kinds of content does your organization produce? List them. For each content type, identify differentiating dimensions, the about-ness. How would you distinguish 1,000 instances of this type? What questions would help you find the right one?
Then define facet values. For each dimension, what are the valid values? Some are controlled vocabularies like product names or department names. Some are dates. Some are free-form tags.
Identify cross-cutting facets next. Some dimensions apply across multiple content types, such as audience, geography, and effective date. These become enterprise facets. Finally, map relationships. What content is related to what other content? Products to documentation. Policies to procedures. Troubleshooting guides to parts lists.
Progressive Enhancement: How to Build Without Boiling the Ocean
The biggest objection to rich content models is effort: "We cannot tag 100,000 documents with 15 metadata fields. That would take years."
You are right, which is why you do not do it all at once. Progressive enhancement starts simple and adds richness over time, leveraging AI and usage data to do most of the work.
Phase 1: Core Metadata (Day 1)
Start with the absolute minimum, 5 fields that any content creator can fill in 2 minutes: content type (is-ness), title, owner, created or modified date, and 2-3 essential tags. This is not rich enough for sophisticated AI retrieval, but it is enough to launch. You can find content. You can identify owners. You can track freshness.
Phase 2: AI-Assisted Enhancement (Week 1-4)
Now let AI do the heavy lifting. AI reads document content and suggests topics, identifies related documents based on similarity, proposes audience and use case based on content signals, and extracts entities like product names, people, and dates. Humans review and approve AI suggestions rather than tagging from scratch. This is 10x faster than manual tagging. After this phase, you have 10-15 metadata fields per document with reasonable accuracy.
Phase 3: Usage-Driven Enhancement (Month 1-3)
User behavior continuously enriches the model. Query analysis reveals what searches find this content, so those terms get added as keywords. Co-access patterns show what content is frequently accessed together, building relationship links. Feedback signals show what content gets positive ratings, boosting its visibility, and what gets negative ratings, flagging it for review. Gap identification reveals what queries return no results, which are content gaps to fill. This phase runs automatically and continuously. The content model gets smarter as people use it.
Phase 4: Continuous Refinement (Ongoing)
The model is never "done." It evolves. When AI retrieves the wrong content, add differentiating metadata to help distinguish it. When users cannot find content, add metadata that matches their search language. When new products or topics emerge, extend the taxonomy. When terminology shifts, update controlled vocabularies.
The timeline: Day 1 is core metadata only, 5 fields, 2 minutes per document. Week 1, AI enriches up to 15 fields. Month 1, usage patterns emerge and relationships auto-populate. Ongoing, continuous refinement based on AI performance.
The Metadata Paradox (And How to Escape It)
Every content architect faces the same tension. Too little metadata means AI cannot find the right content. Without sufficient descriptive information, AI retrieval is imprecise, and users get irrelevant results or miss content entirely. Too much metadata means nobody fills it in. If you require 47 metadata fields, content creators will skip them, enter garbage data, or stop creating content altogether.
The paradox seems unsolvable: you need rich metadata but cannot get people to create it.
The escape: do not require humans to create rich metadata.
The progressive enhancement model resolves the paradox. Humans provide minimal metadata, 5 fields. AI generates enriched metadata, 10 or more additional fields. Usage data generates relationship metadata automatically. Humans curate and correct rather than create from scratch. The result is rich metadata without a crushing burden on content creators.
Five Warning Signs Your Metadata Is Blocking Scale
Even well-designed content models can become scaling blockers. Watch for these warning signs.
Metadata debt shows up when 60% of documents are missing critical fields. AI cannot differentiate content, retrieval accuracy drops, and users learn not to trust results. The fix is to prioritize high-value content first, use AI-assisted backfill for bulk remediation, and accept that 100% coverage is not necessary.
Metadata inconsistency shows up when the same concept is tagged 15 different ways: "remote work" versus "work from home" versus "WFH" versus "telecommuting" versus "distributed work." AI cannot connect related information, and a search for a single term misses content tagged with inconsistent synonyms. The fix is controlled vocabularies with synonym mapping, auto-suggest from taxonomy rather than free-form entry, and periodic normalization passes.
Metadata overload shows up when there are 47 metadata fields, all marked required. Content creators give up, fields get left blank or filled with placeholder text, and data quality plummets. The fix is to ruthlessly reduce required fields to 5-7, make everything else optional or AI-generated, and apply the test: would a new employee get this right?
Metadata abandonment shows up when a beautiful schema is designed, documented, and approved, and nobody uses it. Content is not tagged, cannot be found, and the investment in design is wasted. The fix is to make metadata entry frictionless by embedding it in existing workflows rather than adding steps, and to show immediate value through better search results and auto-surfaced content.
Conflicting metadata shows up when the same document is tagged differently by Sales, Marketing, and Support. AI gives different answers depending on which version it retrieves, and users get inconsistent information. The fix is single source of truth rules, cross-department governance, and authoritative system designation.
The Scaling Test
Here is a simple test for whether your content model will scale: could a new employee tag 10 documents correctly in 20 minutes using only your metadata guidelines?
If yes, your model is simple enough to scale. If no, your model is too complex. Simplify required fields. Improve documentation. Add AI assistance.
This test exposes the gap between designed elegance and operational reality. Many content models look beautiful in documentation but collapse when real humans try to apply them under time pressure.
Building Adaptive Models for GenAI
When designing content models specifically for AI retrieval, five principles matter most.
Multiple discovery paths: every piece of content should be findable through multiple facets. If content can only be found one way, you will miss users who think differently. Design check: for any document, can you identify at least 3 different queries that should retrieve it?
Relationship-first thinking: GenAI excels when it can traverse relationships. Design your model to make relationships explicit, from products to related documentation, procedures to related policies, troubleshooting guides to related parts, and case studies to related products and solutions.
Graceful incompleteness: your model should work even when the metadata is incomplete. Not every document will have every field filled. Design retrieval to use available metadata rather than failing when fields are empty.
Evolution by design: build in mechanisms for the model to evolve, including regular taxonomy review cycles, feedback loops from AI performance, processes for adding new facets when needed, and deprecation paths for obsolete terms.
Human-AI collaboration: design for the hybrid model from the start, with core fields for humans to complete, extended fields for AI to generate, relationship fields for usage data to populate, and all fields for humans to review and correct.
Case Study: E-commerce Retailer
A major e-commerce retailer with 2.3 million products had a rigidly hierarchical taxonomy. Products lived in exactly one category. When they implemented a GenAI shopping assistant, queries like "gifts for outdoorsy dads under $100" failed because the system could not cross-reference product category, recipient persona since "dad" was not a taxonomy term, use case since "outdoorsy" describes activities, and price range, which is metadata rather than hierarchy.
After migrating to a faceted model, gift-related query accuracy improved from 34% to 78%, cross-category recommendations increased by 156%, customer satisfaction with the AI assistant rose from 3.1 to 4.4 out of 5.0, and conversion rate on AI-assisted sessions improved 23%.
The key insight: the AI did not need better algorithms. It needed content described in the dimensions customers actually use.
The Payoff
Organizations that move from rigid hierarchies to adaptive content models see measurable improvements across several dimensions. Retrieval accuracy improves because AI finds relevant content more consistently through multiple discoverable paths. User satisfaction improves because users find what they need faster without navigating a hierarchy that does not match their mental model. Content reuse improves because the same content serves multiple use cases without arbitrary categorical boundaries siloing it. Maintenance efficiency improves because organizational changes do not require taxonomy restructuring, since facets are independent. And AI improvement compounds because rich metadata gives AI more signals for relevance ranking, leading to better results over time.
The shift is not trivial. It requires rethinking how you describe content, investing in AI-assisted enrichment, and building processes for continuous refinement.
But the alternative, forcing multi-dimensional content into rigid hierarchies and watching AI struggle to retrieve it, is no longer viable.
The Bottom Line
Hierarchical taxonomies were designed for a world of physical filing. Digital content, especially content that needs to serve AI retrieval, requires a multi-dimensional description.
Faceted classification describes content across multiple independent dimensions instead of forcing it into branches. Is-ness and about-ness define what content it is and what makes each instance different. Progressive enhancement starts simple, lets AI enrich, lets usage data enhance, and refines continuously. The metadata paradox resolves when AI and usage data do most of the work, not humans.
Your taxonomy is not failing because it is poorly designed. It is failing because hierarchies are the wrong model for AI-powered retrieval.
The organizations that figure this out first will have AI that actually works. Those who keep polishing their hierarchies will keep wondering why their AI cannot find anything.
