This field note is a working method, not a universal scoring model. Use it to organize evidence, expose uncertainty and define the next test.
01. Model relationships before links
Internal links should express how pages relate: parent, child, sibling, prerequisite, alternative or next step. Starting with anchor-text targets produces arbitrary modules. Define the content model and user decision first.
A diagram of page types and permitted relationships often exposes missing hubs and duplicate roles.
02. Measure depth with context
Click depth is useful, but a page three clicks away in a clear specialist path can be stronger than a page linked indiscriminately from every footer. Measure reachability, number of distinct relevant sources and the prominence of links in rendered content.
Separate global chrome from contextual editorial links in reporting.
03. Find orphans using multiple sources
A crawler sees only linked URLs. Compare crawl output with sitemaps, analytics, Search Console, CMS exports and backlink data. Classify orphans: intentionally retired, newly published, campaign-only, broken workflow or genuinely missing from architecture.
Do not bulk-link every orphan; decide whether each page deserves a maintained role.
04. Design modules around decisions
Related articles, next-step guides, category trails and product alternatives serve different decisions. Give each module an explicit selection rule, fallback and maximum count. Manual curation and automation can coexist when ownership is clear.
Validate the actual rendered destination and canonical URL, not only CMS identifiers.
05. Control taxonomy growth
Tags and categories create navigation promises. Require naming rules, minimum viable populations, indexability criteria and retirement procedures. Near-synonymous archives divide signals and confuse editors.
A taxonomy register with owner, purpose and included content types prevents silent proliferation.
06. Validate change by cohort
After revising links, monitor discovery, crawl frequency, depth and search performance for affected pages against controls. Changes in ranking alone cannot prove the linking mechanism.
Check that new modules remain populated and correct after content updates, not just on release day.
Operational worksheet
Use one row per URL pattern or content cohort. Record the symptom, expected behavior, evidence source, conflicting observations, suspected mechanism, population size, owner and next validation date. Preserve examples that do not fit the leading theory; they often reveal a second template or release path.
- Model relationships before links: record evidence, owner and acceptance test.
- Measure depth with context: record evidence, owner and acceptance test.
- Find orphans using multiple sources: record evidence, owner and acceptance test.
- Design modules around decisions: record evidence, owner and acceptance test.
- Control taxonomy growth: record evidence, owner and acceptance test.
- Validate change by cohort: record evidence, owner and acceptance test.
Choose controls before changing the site
Select unaffected URLs that share the same template, age range and demand profile as the affected cohort. Controls make it possible to distinguish a technical recovery from seasonality, a broad ranking update or a reporting change. Record their status before release and inspect them on the same schedule as changed URLs. If controls move in the same direction, reconsider the proposed mechanism before claiming success.
Keep evidence at URL and pattern level
Site-wide totals are useful for orientation but poor for implementation. Attach every finding to an example URL, the rule or template that produced it, and an estimate of the affected population. Store the exact observation date because crawling and indexing evidence changes. When platform reports disagree, preserve both observations and add the next test; do not average contradictory states into a misleading score.
Turn findings into acceptance criteria
An engineering ticket should describe expected user and crawler behavior. Name the response status, rendered content, indexability directive, canonical target, internal-link source and sitemap state where relevant. Include examples that must change and controls that must remain unchanged. “Fix canonical tags” is not testable; “pagination URLs declare self-canonicals while filtered duplicates consolidate to the clean category URL” is.
Sequence validation by processing delay
Some checks are immediate: deployed markup, response headers, links and redirect behavior. Crawl discovery takes longer. Index selection and traffic response can take longer still. Separate these checkpoints so a team does not roll back a correct release because a search platform has not reprocessed the cohort. Equally, do not wait weeks to discover that the production template still emits the old directive.
Record decisions that reject a recommendation
A review is still useful when the team decides not to implement a finding. Record the reason: low reach, weak confidence, unacceptable user impact, platform constraint or higher-priority work. Add a trigger for reconsideration, such as growth beyond a URL threshold or a future migration. This prevents the same issue from being rediscovered without the context behind the original decision.
Close the loop with ownership
Assign one owner for implementation and another, where possible, for validation. Define the release marker, expected observation window and reporting location. A recommendation without ownership becomes an archive; a recommendation with a measurable test becomes an operating change. The final record should say what happened, what remained uncertain and which evidence would justify another iteration.
Document the smallest useful next test
When evidence remains incomplete, avoid turning uncertainty into a broad recommendation. Specify the smallest reversible test that can distinguish competing explanations, the URLs included, the observation period and the condition for stopping. Small tests protect users and engineering time while producing evidence that a later team can understand. Record negative results as carefully as positive ones; they narrow the system boundary and prevent repeated work.
Limits and interpretation
Search platform reports are sampled, delayed and interpreted by systems outside the site owner’s control. A clean technical implementation does not guarantee indexing, traffic or rankings. Treat recommendations as risk-reduction and diagnostic work, then validate changes against stable cohorts and business outcomes.