Running Your First Crawl on swissknifeseo.ai: What the App Actually Shows You
A walk through the first scan on swissknifeseo.ai: what the crawler reports, how issues get ranked by cost rather than count, and where the AI citation...
Read the guide ›
Two pages on the same site can look interchangeable. Same template, same tidy title, similar length, both indexed, both loading quickly. Then you ask an answer engine a question that either one could answer, and only one of them gets named as a source. The other might as well not have been published.
That gap almost never comes down to the writing. It sits in the plumbing: how the page is served, what it declares about itself, and whether anything on your own site points at it. All of that is measurable from a crawl, which is why the AI readiness score in Swiss Knife SEO is assembled from eight weighted signals instead of one number pulled out of the air. Walk the twins through the same eight and the one being ignored usually gives itself away on two of them.
The screenshots below are this site's own report, not a mock up, which is also why one of the numbers in them is embarrassing. A readiness number that says 61 and stops there is a horoscope. The version worth having shows its work, so every signal carries its own percentage, its own weight, and a sentence naming what it measured. Anything scoring under 60 gets pulled out separately as a blocker, because that is the part costing you answers this week rather than in theory.
The weights are uneven on purpose. Access is worth 25 points by itself, since a page a model cannot fetch scores zero on everything that comes after it. Readability carries another 30. Understanding takes 15, and the remainder splits across navigation, semantics, and identity. Twin pages normally sit within a point of each other across most of that stack, then diverge hard on a single line.
Two of the eight cover reach. One checks how many AI crawlers your robots.txt actually permits, weighted at 20. The other simply checks whether a robots.txt exists at all, worth 5, on the reasoning that a missing file is a policy nobody ever wrote down.
This is where twin comparisons get uncomfortable. A rule blocking one directory, a plugin default that shipped a blanket disallow for anything that is not Googlebot, a staging rule that survived launch: any of those can close one section while the rest of the site stays wide open. GPTBot allowed, ClaudeBot allowed, PerplexityBot blocked is a completely ordinary reading, and it usually surprises the person who owns the site.
The report answers it as a table rather than a score, because the only useful form of this question is per bot. Thirteen answer-engine crawlers, each with the vendor behind it and its current status on your robots.txt, plus a copy-ready block for whichever way you want to go.
Two more signals ask whether a machine can see your answer once it arrives. Server-rendered content, weighted 15, measures the share of sampled pages that hold their copy without running any JavaScript. Content fundamentals, also 15, is blunter than that: a title, exactly one H1, and enough body copy to be worth quoting.
Rendering is the classic twin killer. Both pages look identical in a browser. One was built as a static template, the other pulls its main copy from a client-side component, so the raw HTML that arrives first is a shell with a spinner in it. Search crawlers often come back and render it later. Answer engines are far less patient, and several of them never execute JavaScript at all.
Structured data, weighted 15, looks at how many analyzed pages carry valid JSON-LD. Entity identity, worth 10, asks something narrower: does the site define a primary organization or person anywhere a machine can read it. Schema markup is still the plainest way to tell an engine what a page is, and one page carrying it while its twin does not is normal on any site where markup arrived by template rather than by rule.
Identity does more work than it looks like it does. A model choosing between two sources is choosing something to attribute an answer to. A page sitting on a site with a clear publisher entity is easier to credit than one floating on a domain that never says who runs it, which is also part of why the pages that win featured snippets tend to be the same ones that get quoted by assistants.
That is the zero at the top of this page, and it is ours. The homepage here publishes FAQ markup and no organization block, the about page publishes no JSON-LD at all, and the only Organization anywhere on the site sits nested inside each article's publisher field, which is not the same as declaring who runs the place. A machine reading this site can tell what every page is about and still not know whose it is. It went on the fix list the moment the report loaded, which is the honest use of a number like this: not a grade, a queue.
Discoverable key pages, weighted 10, counts indexable pages with at least one inbound internal link that also sit within three clicks of the homepage. It is the cheapest signal on the list to fix, and the one most often failing quietly.
Orphans are accidents, not decisions. A category gets restructured. A nav item is trimmed to make room. A campaign page goes live from a spreadsheet and never gets linked from anywhere else. The page ranks for nothing, gets fetched rarely, and never enters the pool a model draws from, no matter how good the copy is. That is the argument for treating the way you link into your own best content as real work rather than housekeeping.
Accessible semantics carries 10 points and reads the automated accessibility score: names on controls, labels on inputs, a sane heading order, a document structure something can walk end to end. Nobody writes alt text for an answer engine. The overlap is a side effect of the same discipline, because a page built to be understood by software that cannot see it turns out to be easier for every machine to read.
Run both pages through the same report and the shape of the difference is usually dull. One is blocked for a bot the other is not. One survives without JavaScript and the other does not. One has three internal links landing in its body copy while the other has none at all. Rarely is it a subtle judgment about quality, and that is good news, because mechanical faults have mechanical fixes.
The reason a crawl belongs next to the citation data is that neither half explains anything alone. A tracker tells you Gemini did not cite you. The crawl tells you the page renders client side, which is why. We wrote more about that fusion in the background on the platform, and the first crawl walkthrough covers how to read a full report without drowning in it.
Pick the single page you would be most annoyed to see a competitor quoted on instead of you. Check the two access signals first, because they are binary and they invalidate everything else. Then check whether the answer survives with JavaScript switched off. Those two passes take about ten minutes and they resolve most twin mysteries before you touch the markup or the internal links.
After that it turns into ordinary maintenance rather than a project. Watch the blockers list, fix the cheap ones as they surface, and re-check after each change so you can see which move actually earned the citation. The step by step version of each fix lives in the help center, and if you would rather hand the whole thing over, our SEO service page explains what that looks like. Otherwise, tell us which page keeps losing and you will get a straight answer within a working day.
Related reading: what search engine optimization actually covers, the six types of SEO that increase traffic, and how search intent shapes the questions people ask. The rest sits in the SEO archive and on the blog.
Run a free scan in Swiss Knife SEO, the platform we built and use on every client engagement. No account needed for the first audit.
A walk through the first scan on swissknifeseo.ai: what the crawler reports, how issues get ranked by cost rather than count, and where the AI citation...
Read the guide ›The help center is a public reference, not a ticket queue. It carries 35 guides, 50 FAQ answers and 45 glossary terms, and you can read...
Read the guide ›We kept asking questions our tools could not answer, so we wrote our own crawler. It grew into a platform that scores internal links by where...
Read the guide ›