# Strand CMS — full content > An open-source, agent-first publishing system for programmatic blogs and news sites. MDX in Git, SEO and AI-search built in. --- # 8 Best Ghost CMS Alternatives in 2026 (Open-Source & Hosted, Ranked) *By The Strand CMS Team · 2026-08-10 · 12 min read* Canonical: https://www.strandcms.com/blog/best-ghost-cms-alternatives > **Summary:** The right Ghost CMS alternative depends on whether you need memberships and newsletters, a structured content API, or a Git-native publishing workflow. This comparison maps eight options to those jobs. Choosing among Ghost CMS alternatives is not a contest with one universal winner. Ghost is a focused publication product with a polished editor, memberships, newsletters, themes, and APIs ([Ghost documentation](https://ghost.org/docs/)). The alternatives below solve different problems: broader site management, structured content delivery, or a repository-first workflow. This roundup is for a team deciding how to publish a blog or content site in 2026. It evaluates the operating model, strongest use case, trade-offs, and pricing direction from first-party documentation. Pricing changes; check the linked pages before buying. ## Quick ranking: eight Ghost CMS alternatives | Rank | Alternative | Best for | Operating model | Pricing view | |---|---|---|---|---| | 1 | Strand | Agent-first, Git-native programmatic publishing | Open source, MDX in Git | Software is open source; hosting is separate | | 2 | WordPress | Broad websites and extensibility | Open source core plus hosted options | Hosting, services, and extensions vary | | 3 | Sanity | Structured content across multiple front ends | Hosted/headless platform | Check current plan and usage terms | | 4 | Payload | TypeScript teams building an app-backed CMS | Open source framework plus hosting | Infrastructure and team operations vary | | 5 | Strapi | API-first content delivery | Open source with hosted offering | Check edition and hosting costs | | 6 | Plain Markdown in a repository | Small sites with technical owners | Files plus a site generator | Software can be $0; labor and hosting remain | | 7 | Ghost self-hosted | Ghost features with infrastructure ownership | Open source software you operate | Infrastructure and maintenance are yours | | 8 | Contentful | Managed structured content operations | Hosted headless CMS | Check current workspace and API pricing | The ranking is deliberately job-based. Strand ranks first only for a repository-first publishing team; it is not a replacement for Ghost's membership product or a managed editorial suite. ## 1. Strand: best for agent-first publishing [Strand CMS](https://github.com/BowTiedSwan/strand) stores posts as MDX in Git and pairs a validated content schema with generated SEO and AI-search artifacts. Its MCP server gives an agent publishing operations such as creating drafts and validating posts. That makes it a strong Ghost alternative when review happens in Git and the repeatable unit is a content commit. The trade-off is equally important: there is no Ghost-style hosted membership product or CMS UI. Your team owns hosting, editorial policy, and the surrounding application. Start with the [best Ghost CMS alternatives pillar](/blog/best-ghost-cms-alternatives) and [Introducing Strand](/blog/introducing-strand) for the product model. ## 2. WordPress: best for breadth [WordPress documentation](https://wordpress.org/documentation/) covers a general-purpose publishing platform with a large ecosystem. It is the practical choice when a site needs familiar administration, many integrations, and more than a publication workflow. That breadth creates assembly work. Hosting, security, themes, plugins, and updates become part of the operating model. Choose WordPress when compatibility and ecosystem depth matter more than a narrow, agent-first stack. ## 3. Sanity: best for structured content [Sanity's documentation](https://www.sanity.io/docs/) describes a structured-content and API-oriented approach. It fits teams that need one content model delivered to web, product, or other channels and want a managed service around the editing experience. A headless workflow is a different product category from Ghost. It can be a better foundation for multi-channel applications, but it requires a front end and more implementation decisions than an integrated publication. ## 4. Payload: best for TypeScript application teams [Payload's docs](https://payloadcms.com/docs) position it as a TypeScript CMS and app framework. That is attractive when developers want content and application code close together and are prepared to operate the resulting stack. It is less attractive when the requirement is “give editors a hosted blog this afternoon.” Evaluate the framework, hosting, auth, media, and editorial responsibilities together rather than comparing only the admin screen. ## 5. Strapi: best for API-first teams [Strapi's documentation](https://docs.strapi.io/) focuses on API-first content management. It is a reasonable Ghost alternative for teams whose primary requirement is exposing structured content to several clients. The price of flexibility is implementation and operations. You still need a front end, deployment plan, permissions model, and content workflow. That can be exactly right for an application team and excessive for a single publication. ## 6. Plain Markdown: best for the smallest technical site A plain Markdown repository is an honest alternative when a developer owns the site, the content model is simple, and there is no need for memberships or a hosted editorial interface. The software cost can be $0, but hosting, review, design, backups, and maintenance are not magically free. Choose this option if you value minimum moving parts more than built-in schema, MCP operations, or a richer publishing core. ## 7. Ghost self-hosted: best when Ghost is the requirement Ghost's official [documentation](https://ghost.org/docs/) covers the product and deployment surface, while its [hosted pricing page](https://ghost.org/pricing/) is the reference for Ghost(Pro). Self-hosting keeps the Ghost software model but transfers infrastructure and maintenance to the operator. This is not really leaving Ghost; it is choosing a different operating arrangement. It can be the right answer when memberships, newsletters, and Ghost's editorial model are non-negotiable. ## 8. Contentful: best for managed enterprise content Contentful is included as a managed headless option for teams that prioritize central governance and structured delivery. It belongs in the same evaluation as Sanity and Strapi, not as a direct replacement for every Ghost feature. Confirm current plan limits, API usage, environments, and editorial needs from Contentful's own materials before making a cost comparison. A vendor shortlist should not turn an unverified price into a fact. ## How to choose without copying Ghost's feature list - Choose **Ghost** when memberships, newsletters, and a focused publication editor are the center of the business. - Choose **Strand** when posts should be reviewable as Git changes and agents need a validated publishing contract. - Choose **WordPress** when ecosystem compatibility outweighs operational simplicity. - Choose **Sanity, Payload, Strapi, or Contentful** when structured content must serve multiple applications. - Choose **plain Markdown** when the site is small and a developer wants the fewest dependencies. The best Ghost CMS alternative is the one whose operating model matches the team. Do not buy a headless platform to solve a membership problem, or a Git-native system when nontechnical editors need a hosted workspace. ## Why this matters A CMS decision determines where review, cost, and failure handling live. Ghost centralizes much of that for a publication. A repository-first system exposes it as code and checks. Both can be honest choices; the mistake is hiding the trade-off behind an “alternative” label. For the Git-native case, see [Introducing Strand](/blog/introducing-strand). For the direct product comparison, continue to [Strand vs Ghost](/blog/strand-vs-ghost). The supporting articles in this cluster cover the narrower decisions: [Ghost llms.txt](/blog/ghost-llms-txt-analysis), [migrate from Ghost](/blog/migrate-ghost-to-strand), [Ghost Pro pricing](/blog/ghost-pro-pricing-breakdown), and [what agent-first cuts](/blog/what-agent-first-cuts). ## A decision tree for Ghost switchers Start by naming the job Ghost is doing today. If it is primarily a publication with newsletters and memberships, compare hosted Ghost with the cost of keeping those functions elsewhere. If it is mainly a blog that developers already edit in Git, compare Strand and plain Markdown before adding a larger platform. If the same records must power a product, app, and website, evaluate a headless CMS such as Sanity, Strapi, Payload, or Contentful. Next, name the operator. A technical founder comfortable with repositories can accept responsibilities that a marketing team should not have to absorb. Conversely, a team that has already built CI and review around code may find a GUI-first workflow slower for repetitive programmatic publishing. The right answer follows the people who will run it, not the number of checkboxes in a feature grid. Finally, test the difficult content, not just a fresh post. Import a post with images, embeds, tables, redirects, and metadata. Ask who fixes it when the conversion fails. That small exercise exposes more than a demo because it makes the switching cost visible. ## Evaluate the migration, not just the demo A Ghost alternative can look excellent in a feature matrix and still be wrong for a real publication. Before switching, run a small proof with ten representative posts. Include one long article, one post with images, one with an embed, one with a table, one with a custom HTML block, and one with a redirect. Measure whether the content remains readable, whether metadata survives, and whether a reviewer can correct the output without opening the source code. Then document the answers to operational questions. Who owns the repository? Who can approve a release? What happens when a source URL dies? Where do uploads live? How are previews generated? How are member records handled? How do you restore yesterday's version? A tool that cannot answer these questions is not necessarily bad, but its missing answers are part of the switching cost. ## A fair shortlist by team shape For a solo technical founder, plain Markdown or Strand may be enough. Plain Markdown has the smallest surface area; Strand is more opinionated and adds schema, generated artifacts, and agent operations. The founder should choose the extra machinery only when those checks and outputs earn their keep. For a publication with a membership business, Ghost deserves to remain on the shortlist. Its strengths are not a footnote: the integrated publication model, memberships, newsletters, and themes can be the reason to stay. A Ghost alternative that removes those functions without a replacement is not a like-for-like recommendation. For a product team with multiple clients, Sanity, Strapi, Payload, or Contentful may fit better than either Ghost or Strand. Their value is structured delivery and application integration. The trade-off is building and operating a front end and a broader content system. For a general marketing site with many integrations, WordPress remains a reasonable candidate. Its ecosystem can solve requirements that a narrower system intentionally does not. The team should budget for plugin selection, updates, security, and governance rather than assuming breadth is free. A buyer should also distinguish “alternative” from “replacement.” A replacement is expected to preserve the critical job and migration path. An alternative may deliberately change the job: a developer might trade Ghost's editor for Markdown, or a product team might trade an integrated blog for a structured API. State that change plainly so the shortlist remains useful. ## A practical selection scorecard Score each candidate from 1 to 5 against your actual constraints, and write evidence beside every score: | Question | Why it matters | |---|---| | Can the team publish the needed content type? | Prevents a blog tool from being selected for an application job. | | Can the team review changes in its preferred way? | Aligns the CMS with editorial and engineering habits. | | Are important integrations native or easy to maintain? | Avoids replacing one manual workflow with three. | | Can the system preserve URLs and metadata? | Protects existing readers and search equity during a switch. | | Is the total cost visible? | Includes hosting, labor, email, and maintenance. | | Does the system fail safely? | Makes bad drafts and broken releases easier to catch. | Do not award a high score for a feature you will not use. Do award a high score for a constraint you cannot compromise, such as membership billing or repository review. A second useful test is the exit test. Ask how difficult it would be to leave after two years. Can you export clean content? Are URLs documented? Are images stored in a portable form? Does your team understand the schema, or is the publication trapped in custom behavior that only one vendor or contractor can explain? Portability has value even if you never exercise it. Also test the review surface with a deliberately flawed draft. Put a missing source, a broken internal link, an overlong title, and a malformed image reference into the sample. Observe whether the system rejects the problems, shows them clearly, or lets them reach production. This is where a content team's quality policy becomes a property of the tool rather than a line in a document. ## Questions to ask vendors and maintainers Ask how content leaves the system. Is the output server-rendered, available as Markdown, delivered through an API, or dependent on a browser runtime? Ask how a title, slug, author, date, image, and source are validated. Ask whether a failed build blocks publication or merely produces a warning. Ask about the boring paths too. How do you export your content? How do you restore a previous version? Can you move images without losing URLs? What happens if a third-party integration stops working? Does the service offer a documented migration path? These questions matter more after the purchase than a polished demo does before it. For AI-assisted publishing, ask what is actually automated. “AI-ready” can mean a Markdown endpoint, a prompt integration, a content API, or simply a marketing label. Require a concrete URL, repository workflow, or documented tool before treating the capability as part of the product. ## A note on hosted versus open source Hosted and open-source choices are not opposites. Ghost has an open-source repository and a hosted Ghost(Pro) service. A team can choose Ghost's product and pay for operations to be managed, or run software itself and accept the additional work. The same distinction applies to many headless CMS options and to Strand. Make the decision with a responsibility table. Put hosting, updates, backups, security, email, analytics, media, and editorial review in rows. Put a named owner beside each row. If a row has no owner, the system is not ready for production regardless of its license. ## FAQ ### What is the best Ghost CMS alternative for memberships? Ghost remains a strong fit for memberships and newsletters. WordPress is the broader alternative when you need its ecosystem and are willing to assemble the stack. ### What is the best open-source Ghost alternative for Git-based publishing? Plain Markdown is the smallest option. Strand adds a validated MDX publishing core, generated artifacts, and agent tooling; hosting and editorial operations still remain with you. ### Is Ghost better than a headless CMS? For an integrated publication, often yes. A headless CMS is better when content must serve multiple front ends and the team can own the added implementation. ## FAQ **What is the best Ghost CMS alternative for memberships?** Ghost itself remains a strong fit for memberships and newsletters. WordPress is the broader alternative when you need its large extension ecosystem and are willing to assemble the stack. **What is the best open-source Ghost alternative for Git-based publishing?** A plain Markdown repository is the simplest option; Strand adds a validated MDX publishing core, generated SEO and AI-search artifacts, and agent tooling. **Is Ghost better than a headless CMS?** Ghost is usually simpler for an integrated publication. A headless CMS is a better fit when one content model must serve several front ends or applications. ## Sources - [Ghost documentation](https://ghost.org/docs/) — Ghost - [Ghost(Pro) pricing](https://ghost.org/pricing/) — Ghost - [WordPress documentation](https://wordpress.org/documentation/) — WordPress - [Sanity documentation](https://www.sanity.io/docs/) — Sanity - [Strapi documentation](https://docs.strapi.io/) — Strapi - [Payload documentation](https://payloadcms.com/docs) — Payload - [Strand CMS repository](https://github.com/BowTiedSwan/strand) — Strand CMS --- # Ghost llms.txt: Is That Enough for AI Search? (2026 Analysis) *By The Strand CMS Team · 2026-08-10 · 6 min read* Canonical: https://www.strandcms.com/blog/ghost-llms-txt-analysis > **Summary:** Ghost llms.txt is a navigation aid, not a ranking guarantee. Publishers still need useful, crawlable, well-structured pages, accurate sources, and normal technical SEO for AI search visibility. Ghost llms.txt support is useful only if you describe it accurately, because the file is an index for language-model tools rather than a switch that changes how a Ghost site ranks in search results. The [llms.txt proposal](https://llmstxt.org/) presents the file as a curated Markdown guide to a website for language-model use. That makes it a navigation aid. It is not a contract that Google, ChatGPT, or another model must crawl, rank, or cite your pages. The practical answer for a Ghost publisher is therefore “helpful, but not enough.” The Ghost llms.txt question should end in a checklist, not a promise: what does the file list, are the destinations current, and can a reader use those pages without the index? AI search still depends on the underlying pages being accessible, useful, clear, and trustworthy. Google's [AI features guidance](https://developers.google.com/search/docs/fundamentals/ai-optimization-guide) keeps the emphasis on ordinary search fundamentals rather than a magic file. ## Ghost llms.txt: what the file is for A site can be difficult for a language model to understand when important material is spread across navigation, JavaScript, archive pages, and repeated UI. A concise Markdown index can point to canonical explanations, documentation, and other high-value resources. That is the narrow benefit: better orientation. The proposal is not a claim that the file changes the truth of the linked pages or bypasses a crawler's access rules. Keep descriptions factual and links stable. ## What Ghost publishers should verify Ghost's [official documentation](https://ghost.org/docs/) describes its publishing and theme capabilities. Before treating a Ghost-generated or manually added `llms.txt` as a finished AI-search strategy, check four things: 1. **The file is reachable.** Request `/llms.txt` over HTTPS and confirm it returns the intended content. 2. **The links are canonical.** Point to the public pages you want a reader or model to understand, not duplicate paths. 3. **The pages stand alone.** Each target should explain its subject without requiring a client-side interaction to reveal the answer. 4. **The index is maintained.** Remove stale pages and update descriptions when the site structure changes. These checks improve the file's usefulness; they do not establish a guaranteed visibility outcome. ## Why llms.txt is not a ranking switch A ranking promise would require evidence that the relevant search engine uses the file in that way. The cited proposal describes a format, while Google's guidance for AI features points site owners toward helpful content and established search practices. Neither source supports “add this file and rankings rise.” The same caution applies to AI citations. A model may use a page because it is relevant, accessible, and credible, but an index file alone cannot manufacture those properties. Avoid reporting a correlation as causation and avoid claiming a Ghost implementation changes Google rankings without a source. ## The rest of a Ghost AI-search checklist - Write an answer-first introduction that names the page's subject. - Use descriptive headings and short, self-contained sections. - Keep important facts in the rendered HTML and accessible Markdown, not only in interactive widgets. - Link related explanations with descriptive anchor text. - Add sources for statistics, prices, product claims, and other checkable facts. - Review titles, descriptions, canonical URLs, robots directives, and sitemaps through normal technical SEO checks. The [Strand CMS repository](https://github.com/BowTiedSwan/strand) is an example of a different operating model: its content core describes generated machine-readable outputs and validation. That is a product design choice, not evidence that one format guarantees citations. ## Why this matters The phrase “Ghost llms.txt” can make a file sound like the whole strategy. It is not. Use `llms.txt` to make a site's important material easier to navigate, then earn visibility with pages that are useful enough to be cited. For the broader vendor decision, read [best Ghost CMS alternatives](/blog/best-ghost-cms-alternatives). If you are evaluating a Git-native alternative, see [Introducing Strand](/blog/introducing-strand). ## What the file cannot do An `llms.txt` file cannot make a private page public, repair a blocked resource, replace a sitemap, or supply evidence that the page itself does not contain. It cannot turn a thin product page into a useful explanation. It cannot make an inaccurate description trustworthy. It also cannot settle ownership. If a page is sourced from a company announcement, a research paper, or a public statistic, cite that source on the page. A model should not have to infer the authority of a claim from an index entry. ## Ghost-specific implementation questions When reviewing Ghost llms.txt support, separate three questions: - **Generation:** Does the site create the file automatically, or does an operator maintain it? - **Content:** Does it list the canonical pages that matter, with concise accurate descriptions? - **Delivery:** Does the public domain serve the file at the expected path without an accidental redirect or access restriction? The proposal at [llmstxt.org](https://llmstxt.org/) helps define the first two concepts, but the actual Ghost configuration and hosting behavior must be checked on the site being operated. Do not infer a universal implementation from a marketing phrase. ## How to measure usefulness without overclaiming Keep a dated copy of the file and record which pages it links to. Ask a reviewer who did not build the site to find the canonical answer to three real questions using the file and the linked pages. If the paths are confusing or the target pages are weak, improve those pages first. For search performance, monitor ordinary evidence such as crawl errors, impressions, clicks, and citations where a platform exposes them. Do not label a change caused by `llms.txt` unless the measurement design can actually isolate it. In most small sites, many SEO changes happen at once. ## A maintenance routine for Ghost llms.txt Review the file whenever the site's information architecture changes. Remove links to retired pages, update descriptions after a major rewrite, and keep the list short enough to be useful. Ask whether each entry helps a reader locate a canonical answer or merely repeats navigation that already exists. Treat the file as documentation with an owner and a review date. That simple discipline is more defensible than promising that a new format will produce a specific search outcome. A publisher can also use the review to improve the linked pages themselves. If a model or a human cannot understand a target from its title, headings, and opening paragraphs, changing the index description only hides the weakness. Rewrite the page, add the missing evidence, and then update the index. The file should reflect a useful site rather than stand in for one. Keep the maintenance rule simple: every new canonical guide gets considered for the index, every retired guide is removed, and every description is checked for accuracy. That keeps the file aligned with the site instead of turning it into a second, stale navigation system. The file works best as a maintained map, not as a campaign claim. ## FAQ ### What does Ghost llms.txt do? It can provide a curated, machine-readable map of important resources. It does not guarantee that a model crawls or cites those resources. ### Does llms.txt improve Google rankings? The sources here do not support a ranking guarantee. Google's guidance emphasizes useful content and normal search fundamentals. ### Is llms.txt enough for AI search visibility? No. It is one navigation aid alongside accessible pages, clear structure, accurate sourcing, internal links, and technical SEO. ## FAQ **What does Ghost llms.txt do?** It can provide a machine-readable list of important site resources. Treat it as navigation context, not proof that a model will crawl or cite every page. **Does llms.txt improve Google rankings?** There is no basis in the cited sources for promising a ranking boost. Google's guidance emphasizes useful, crawlable content and standard search fundamentals. **Is llms.txt enough for AI search visibility?** No. It should sit alongside accessible pages, clear structure, accurate claims, internal links, and appropriate technical SEO. ## Sources - [The llms.txt proposal](https://llmstxt.org/) — llmstxt.org - [Google guidance on AI features and your website](https://developers.google.com/search/docs/fundamentals/ai-optimization-guide) — Google Search Central - [Ghost documentation](https://ghost.org/docs/) — Ghost - [Ghost(Pro) pricing](https://ghost.org/pricing/) — Ghost - [Strand CMS repository](https://github.com/BowTiedSwan/strand) — Strand CMS --- # Ghost Pro Pricing in 2026: What You Pay For (And the $0 Alternative) *By The Strand CMS Team · 2026-08-10 · 6 min read* Canonical: https://www.strandcms.com/blog/ghost-pro-pricing-breakdown > **Summary:** Ghost(Pro) is a hosted subscription priced on Ghost's official plan page. Self-hosting can remove the software subscription but not infrastructure or labor; Strand's open-source software has the same cost-of-ownership caveat. Ghost Pro pricing is a hosted-service question, so the source of truth is [Ghost's official pricing page](https://ghost.org/pricing/), checked for this article on 2026-08-10. The page can change; use it for the current plan names, prices, and limits rather than treating this explanation as a permanent rate card. The useful comparison is not “paid versus free.” It is **hosted subscription versus operator-owned total cost**. Ghost(Pro) charges for a managed Ghost publication. Self-hosted Ghost and open-source alternatives can remove a software subscription while leaving infrastructure, maintenance, email, backups, and labor on your side. ## Ghost Pro pricing: what the subscription covers The phrase **Ghost Pro pricing** refers to Ghost's hosted offering, not the cost of every possible Ghost deployment. Keep the hosted plan, self-hosted software, and a different open-source system in separate columns. That prevents a $0 software claim from being mistaken for a $0 operating budget. ## What Ghost(Pro) pricing buys Ghost's [documentation](https://ghost.org/docs/) describes a publication product with its own editor, themes, memberships, newsletters, and APIs. Ghost(Pro) packages that product as a hosted service. The exact plan boundary matters: check the official page for current staff, audience, feature, and usage limits before choosing a tier. A hosted plan can be worth paying for when the team values a vendor-operated application and wants to spend less time on deployments, updates, backups, and server troubleshooting. That convenience is part of the cost, not an incidental add-on. ## The current price is a snapshot, not a promise Pricing pages are mutable. Do not copy a monthly number into a long-lived budget without recording the date, billing interval, currency, included limits, and overage rules shown on the page. If a team grows its audience or changes its publication needs, the relevant plan may change too. This article intentionally avoids inventing a price from memory. For an actual purchase decision, open [Ghost's pricing page](https://ghost.org/pricing/) and capture the plan details you will use. ## What self-hosting changes Ghost's [official repository](https://github.com/TryGhost/Ghost) is the source for its open-source code. Running that software yourself can change the subscription line, but it does not make the service literally free. You still need a server, storage, a domain, TLS, monitoring, backups, upgrades, and someone responsible for incidents. Email delivery and membership operations can add their own vendor costs. The amount depends on traffic, audience, architecture, and the operator's time. A fair self-hosting budget lists those line items instead of writing “$0.” ## What the $0 alternative means for Strand The [Strand repository](https://github.com/BowTiedSwan/strand) describes an open-source, Git-native publishing system with MDX content, validation, and generated site artifacts. That supports a **$0 software subscription** framing—not a $0 production bill. Choose Strand when your team wants the repository to be the source of truth and can operate the site, CI, hosting, and editorial checks. If the team needs Ghost's membership and newsletter product or a hosted editor, Strand is not a price-equivalent replacement just because its code is open source. ## Compare total cost of ownership | Cost area | Ghost(Pro) | Self-hosted Ghost | Strand | |---|---|---|---| | Software subscription | Use current Ghost plan page | No Ghost software subscription | No Strand software subscription | | Hosting | Packaged by the hosted service | Operator pays | Operator pays | | Updates and backups | Vendor-managed within service scope | Operator-managed | Operator-managed | | Editorial UI | Ghost product | Ghost product | Git/MDX workflow | | Memberships/newsletters | Ghost product strengths | You operate the stack | Not a parity claim | | CI and validation | Surrounding workflow varies | Surrounding workflow varies | Core repository workflow | The table is a responsibility map, not a promise of identical features or costs. Add your team's labor rate and required services before making a switch. ## Why this matters Ghost Pro pricing is easy to compare as a number and harder to compare as an operating contract. Pay for hosting when it removes work you do not want. Choose open source when you want control and can honestly own the work that comes with it. For the broader alternatives decision, read [best Ghost CMS alternatives](/blog/best-ghost-cms-alternatives). For workflow trade-offs, see [Strand vs Ghost](/blog/strand-vs-ghost). ## Ask what the price includes A hosted plan is not merely a server rental. It may bundle application updates, backups, security work, support, and a managed publishing environment. The useful question is which responsibilities Ghost handles within the plan and which still belong to the customer. Read the plan page for limits that affect your publication: audience size, staff access, newsletters, storage, integrations, and any usage-based conditions. Do not assume every feature in Ghost's general documentation is included in every hosted tier. ## A simple budget model Build three scenarios rather than one headline number: 1. **Hosted:** Ghost(Pro) subscription plus any external services the publication needs. 2. **Self-hosted Ghost:** infrastructure, email, storage, backups, monitoring, upgrades, and operator time. 3. **Open-source alternative:** the same infrastructure categories, plus any missing product capabilities you must build or buy. Use your own traffic and team rates. A technically skilled operator may value control highly, while a small editorial team may rationally pay for a managed service. The cheapest invoice is not always the lowest total cost. ## When the $0 software alternative is a fit An open-source system is a sensible $0-software option when the team already has deployment skills, does not need Ghost's membership product, and values control over the publishing source. Strand's Git/MDX model is aimed at that kind of operator. It is a poor fit if “free” is being used to hide a requirement for a visual editor, hosted email, member billing, or hands-off operations. Price is only one dimension of product fit. ## Pricing questions that change the answer Ask whether your team needs one site or several, whether media storage is substantial, whether email is transactional or newsletter-based, and whether someone is on call for upgrades. Ask how much time a publishing outage costs. Those answers can make a hosted Ghost plan economically sensible even when an open-source license is available. Also separate fixed and variable costs. A domain and small server may be predictable; audience growth, image delivery, email volume, support, and contractor work may not be. Put assumptions and dates beside each estimate. This avoids presenting a temporary pricing-page snapshot as a total-cost forecast. A price comparison should include the cost of switching. Content transformation, redirect testing, design work, staff training, and a period of dual operation all consume time. They may be worthwhile, but they are not zero simply because the destination software has an open-source license. The final question is what kind of uncertainty the team prefers. A hosted Ghost plan makes the subscription line visible and reduces some infrastructure responsibility. Self-hosting or using Strand can increase control and reduce software subscription exposure, while making more operational work explicit. Neither is universally cheaper until the team supplies its own traffic, labor, and reliability assumptions. ## FAQ ### Where is current Ghost Pro pricing? Ghost's official pricing page is the current source. Check it at purchase time because plans and prices can change. ### Is self-hosted Ghost free? The software is open source, but infrastructure, email, backups, maintenance, and operator time remain costs. ### Is Strand a $0 CMS? Its software can be used without a subscription. A production site still needs hosting, a domain, CI, and labor. ### What does Ghost Pro pay for? It is a hosted service around Ghost's publication product. Use Ghost's plan page for the exact current inclusions and limits. ## FAQ **Where can I find current Ghost Pro pricing?** Use Ghost's official pricing page. Prices and plan boundaries can change, so this article does not treat a snapshot as permanent. **Is self-hosted Ghost free?** The Ghost software is open source, but servers, storage, email, backups, maintenance, and operator time still have costs. **Is Strand really a $0 CMS?** The open-source software can be used without a software subscription, but a production Strand site still has hosting, domain, CI, and labor costs. **What does Ghost Pro pay for?** Ghost(Pro) packages hosted operation around Ghost's publication product; use the official plan page for the exact included limits and current price. ## Sources - [Ghost(Pro) pricing](https://ghost.org/pricing/) — Ghost - [Ghost documentation](https://ghost.org/docs/) — Ghost - [Ghost open-source repository](https://github.com/TryGhost/Ghost) — Ghost - [Strand CMS repository](https://github.com/BowTiedSwan/strand) — Strand CMS --- # How to Migrate from Ghost to Strand in an Afternoon (Step-by-Step) *By The Strand CMS Team · 2026-08-10 · 9 min read* Canonical: https://www.strandcms.com/blog/migrate-ghost-to-strand > **Summary:** Migrating from Ghost to Strand is a content transformation project: export and inventory Ghost content, map fields into MDX, move media, recreate redirects, validate every post, and cut over only after review. To migrate from Ghost to Strand in an afternoon, treat the work as a controlled content transformation—not a theme copy. Ghost's [migration documentation](https://ghost.org/docs/migration/) and [Content API documentation](https://ghost.org/docs/content-api/) provide the source-side context; Strand's [repository](https://github.com/BowTiedSwan/strand) provides the destination model of MDX files in Git with validation. A realistic migration has six deliverables: posts, frontmatter, media, URLs and redirects, editorial features, and a cutover plan. Posts are the part that can be scripted most directly. Memberships, subscriptions, themes, and integrations are product decisions and must not be silently discarded. ## Migrate from Ghost: the afternoon plan The phrase **migrate from Ghost** hides several different moves. This guide covers moving published content into Strand's MDX-in-Git workflow; it does not promise that Ghost memberships, newsletters, themes, or integrations have a one-click equivalent. ## Step 1: Freeze the source and make an inventory Choose a short content freeze window. Record the Ghost site's post count, authors, tags, canonical URLs, publication states, image URLs, and any custom HTML or embeds. Also list the things that are not ordinary posts: - Members and subscription tiers - Newsletters and email templates - Theme behavior and custom routes - Forms, analytics, search, and integrations - Redirect rules and external backlinks you care about This inventory is the acceptance checklist. It prevents a migration script from reporting success while losing the parts that made the site work. ## Step 2: Export Ghost content using a documented path Start with Ghost's [migration guidance](https://ghost.org/docs/migration/). Where an integration or API is appropriate, use Ghost's documented [custom integration guidance](https://ghost.org/help/creating-a-custom-integration/) and [Content API](https://ghost.org/docs/content-api/) rather than scraping the rendered theme. Save the raw export unchanged. Keep it outside the transformed output so you can rerun the conversion and compare results. Record the export timestamp and the source version or endpoint used. Do not assume that an export includes every member record, email setting, theme asset, or integration. Mark each category as exported, rebuilt, replaced, or intentionally retired. ## Step 3: Map Ghost fields to Strand frontmatter Create a mapping before writing files. A typical post map looks like this: | Ghost concept | Strand destination | Review question | |---|---|---| | Title | `title` | Is it within the destination schema limit? | | Slug | `slug` and filename | Does the public URL remain stable? | | Excerpt or description | `description` and `summary` | Is it a useful answer, not a truncated sentence? | | Published date | `publishedAt` | Is the timezone preserved? | | Author | `author` | Does the destination author file exist? | | Tags | `tags` | Are names normalized without changing meaning? | | HTML/Markdown body | MDX body | Are embeds, tables, and links valid? | Strand requires more than moving text. Add the destination content type, primary keyword, keywords, FAQ entries, and visible sources required by the site's content contract. Keep the generated MDX readable so a reviewer can inspect the result. ## Step 4: Transform bodies into MDX and move media Convert each body into MDX conservatively. Preserve headings, lists, links, code, tables, and meaningful emphasis. Replace Ghost-specific cards, embeds, and theme shortcodes with supported components or explicit fallback text. Download or copy media into the destination's expected asset location, then update references. Check alt text and image dimensions where available. Open representative posts containing galleries, videos, code, and custom HTML; these are where automated conversions most often look plausible but fail in the browser. Every migrated body should start with ``, not an in-body H1. The template renders the title from frontmatter. Link the post to the [best Ghost CMS alternatives pillar](/blog/best-ghost-cms-alternatives) and to [Introducing Strand](/blog/introducing-strand) as the product context. ## Step 5: Preserve URLs and build redirects Keep the old slug where possible. Build a source-to-destination redirect table for changed paths, including trailing-slash variants and renamed posts. Test redirects with real HTTP requests after deployment. Check canonical tags and internal links so the new site does not point readers back to retired URLs. Do not call a migration complete because the home page works. Crawl the destination list, sample old URLs, and inspect 404s, images, feeds, sitemap entries, and Markdown equivalents. ## Step 6: Validate, review, and cut over Run the destination's schema validation and content QA before the cutover. Fix missing authors, invalid slugs, broken MDX, unsupported components, absent sources, and malformed FAQ data. Then manually review a sample from every post type and every high-traffic URL. Use a two-person sign-off if the site has meaningful traffic: one reviewer checks content fidelity and one checks URLs, metadata, media, and redirects. Keep the Ghost export and redirect table as rollback material. Cut over only after the acceptance checklist is green. ## What does not port automatically A Ghost-to-Strand migration does not promise feature parity. Ghost's docs cover a product with themes, members, publishing, and APIs. Strand's repository describes a Git/MDX publishing core. Plan explicitly for: - Members, paid subscriptions, and email delivery - Theme templates and custom design behavior - Forms, analytics, search, and integrations - Scheduled posts and editorial permissions - Cards or embeds with no destination equivalent The correct outcome may be to keep a separate service for one of these functions or to postpone the migration. A smaller, verified move is safer than a complete-looking import with hidden losses. ## Why this matters Migration is a URL and data-integrity project before it is a CMS preference. The safest path is export → map → transform → validate → review → redirect → cut over. That sequence makes omissions visible and gives you a rollback point. ## Step 7: Run a content fidelity review Automation can tell you that six files parse. It cannot tell you whether a paragraph changed meaning, whether a callout became an empty box, or whether an image now appears beside the wrong caption. Compare the source and destination for title, slug, author, date, body, links, images, and embedded media. Use a review matrix with one row per migrated post. Mark each field as exact, transformed, manually rebuilt, or intentionally omitted. Require an explanation for every omission. This is especially important for custom HTML and cards: a parser may preserve markup that the destination renderer does not support. Review the posts that matter most first. Start with pages receiving organic traffic, pages with external links, evergreen guides, and posts referenced from newsletters. Then sample the long tail. A migration is not safe because the five newest posts look correct. ## Step 8: Plan the cutover window Choose a cutover time that leaves room to test without rushing. Lower the chance of a race by freezing edits in Ghost, taking one final export, transforming the delta, and recording the exact destination commit. Keep the old site available until redirects and representative pages have been checked. At cutover, verify the home page, a post, a tag or archive page, the feed, the sitemap, robots directives, images, and several old URLs. Check from outside the local development environment. A local success does not prove the production host has the same routes or environment variables. ## Rollback and post-launch checks Keep the raw export, the mapping specification, the redirect table, and the last known Ghost deployment details. If the destination has a serious defect, rollback should be a documented operation rather than a stressful reconstruction. For the first days after launch, watch 404s, redirect chains, missing media, canonical mismatches, and reader reports. Fix the highest-impact problems first and record corrections in Git. The goal is not to claim a perfect automated migration; it is to make every remaining difference known and recoverable. A useful migration log records more than errors. Note which transforms were applied to HTML, which embeds were replaced, which images could not be fetched, and which posts needed an editor's rewrite. This makes future migrations easier and gives the team an honest explanation of why the destination does not match the source byte for byte. If the site has comments or member-only content, define the boundary before importing anything. A public post may be safe to transform into a public MDX file; private records may require a different system, a retention decision, or no migration at all. Do not place sensitive data in a repository merely because the export made it easy to access. ## Migration acceptance checklist Before declaring the Ghost migration complete, check each layer separately. At the content layer, compare counts, slugs, titles, dates, authors, tags, and representative bodies. At the asset layer, request every referenced image and inspect alt text, captions, and responsive behavior. At the discovery layer, inspect canonical URLs, the sitemap, feeds, robots directives, and internal links. At the business layer, confirm what happened to members, paid subscriptions, newsletters, forms, analytics, and integrations. A post migration can be technically green while the business is missing a critical audience workflow. Give each item a named owner and a status of rebuilt, replaced, retained elsewhere, or retired. At the release layer, confirm the validator ran against the final commit, the redirects were tested from the public host, and the team knows how to roll back. Keep a short record of the cutover timestamp and the unresolved differences, even when the unresolved list is empty. ## A migration is a project, not an import button The temptation is to start with a converter and discover the requirements afterward. Reverse that order. Write the inventory and mapping first, choose a representative sample, then automate only the transformations you can inspect. Keep a human review step for every content feature that has no exact destination equivalent. That approach may leave a few Ghost features outside the first release. It is better to launch a smaller, verified publication than to claim that members, redirects, media, and custom embeds survived when they did not. The migration is successful when readers can find the right pages and the team knows who owns the new system. Keep the source export read-only during the acceptance period. If a reviewer discovers a missing paragraph or wrong image, correct the mapping or source data and rerun the transformation instead of editing only the destination by hand. That keeps the migration reproducible and reduces the chance that a later rerun silently removes the fix. Record the final redirect tests and the reviewer who approved the cutover. This small paper trail turns a risky switch into an auditable release and gives the team a clear starting point for later corrections. A documented handoff is part of the migration deliverable. The receiving team should have the export location, mapping notes, redirect table, validation command, and rollback contact before launch day. Without that handoff, a technically successful import can still become an operational failure. ## FAQ ### Can every Ghost feature migrate automatically? No. Posts and metadata are the easiest part. Themes, members, subscriptions, and integrations need separate rebuild or replacement decisions. ### What should I export first? Export documented content and metadata, then inventory media, redirects, members, themes, and integrations separately. ### Will URLs stay the same? They can when slugs and routes are preserved, but test redirects and canonical URLs rather than assuming it. ### How do I validate the migration? Run the destination schema validator and batch QA, then manually review rendered content, links, images, metadata, and representative edge cases. ## FAQ **Can I migrate every Ghost feature to Strand automatically?** No. Posts and metadata can be transformed, but themes, members, subscriptions, and integrations need separate decisions and manual review. **What should I export before leaving Ghost?** Export the content and metadata available through Ghost's documented migration and admin/integration paths, then separately inventory media, redirects, members, themes, and integrations. **Will Ghost URLs stay the same after migration?** They can if you preserve slugs and configure the destination routes, but verify every redirect and canonical URL rather than assuming a one-to-one mapping. **How do I validate a Ghost-to-Strand migration?** Run the Strand schema validator and batch QA, then review rendered posts, links, images, metadata, redirects, and representative content manually. ## Sources - [Ghost migration documentation](https://ghost.org/docs/migration/) — Ghost - [Ghost custom integration guidance](https://ghost.org/help/creating-a-custom-integration/) — Ghost - [Ghost Content API documentation](https://ghost.org/docs/content-api/) — Ghost - [Ghost documentation](https://ghost.org/docs/) — Ghost - [Strand CMS repository](https://github.com/BowTiedSwan/strand) — Strand CMS --- # Strand vs Ghost for Programmatic Publishing in 2026 (Which Wins?) *By The Strand CMS Team · 2026-08-10 · 11 min read* Canonical: https://www.strandcms.com/blog/strand-vs-ghost > **Summary:** Strand wins when programmatic publishing means validated MDX commits and agent-operated workflows. Ghost wins for a polished human editor, memberships, newsletters, and an integrated publication product. **Short answer:** Strand wins the Strand vs Ghost decision when programmatic publishing means agents creating MDX, deterministic validation, and Git review. Ghost wins when the core job is human-led publication with memberships, newsletters, themes, and a mature integrated product. This is a workflow comparison, not a claim that one product wins every category. Both products can publish a content site. They put control in different places: Strand puts the source of truth in Git; Ghost puts the publication experience in its CMS. Ghost also documents a [Content API](https://ghost.org/docs/content-api/) for reading content into another application. ## Verdict table | Factor | Strand | Ghost | Winner | |---|---|---|---| | Agent-operated drafting | MCP operations and MDX files in Git are documented in the [Strand repository](https://github.com/BowTiedSwan/strand) | APIs and integrations support automation around Ghost | Strand for a Git-first agent workflow | | Human editing | No CMS UI is the intended model | Polished editor and publication workflow are core Ghost strengths ([docs](https://ghost.org/docs/)) | Ghost | | Validation | Schema and repository checks can block an invalid change | Validation depends on Ghost's product and the surrounding workflow | Strand for code-review gates | | Memberships and newsletters | Not a native parity claim | Native publication strengths documented by Ghost | Ghost | | Multi-front-end delivery | Generated site output and repository ownership | Content API supports read delivery to other clients | Depends on architecture | | Software license | Open-source repository | Open-source Ghost code exists; Ghost(Pro) is hosted and priced separately | Depends on operation | | Ownership of operations | Team owns hosting, Git, CI, and application | Ghost(Pro) centralizes hosted operations; self-hosting transfers them to you | Ghost(Pro) for less infrastructure work | ## Strand vs Ghost: workflow verdict The primary keyword describes the decision precisely: **Strand vs Ghost** is a comparison of publishing workflows, not a generic CMS scorecard. Compare the source of truth, review gate, publishing surface, and operational burden before comparing feature names. ## Workflow: commit review versus editor review A Strand workflow starts with an agent or writer changing an MDX file. The change can be inspected in Git, checked against a schema, and rejected before publication if it fails deterministic gates. That is useful for a programmatic content team with code review habits. Ghost starts from its editor and publication model. That is a genuine advantage for human editors who need a visual, integrated workspace. Automation can call APIs or build integrations, but the central review object is not the same Git commit. ### What Strand does better - Treats content as versioned files that fit normal Git review. - Gives an agent a narrow content contract rather than an unconstrained editor surface. - Makes validation and the generated output part of the publishing workflow. ### What Ghost does better - Provides a mature human publishing experience. - Includes product strengths around memberships, newsletters, and publication themes. - Offers a hosted Ghost(Pro) path for teams that do not want to run the application themselves. ## Artifacts and delivery The [Strand repository](https://github.com/BowTiedSwan/strand) describes MDX-in-Git and generated SEO and AI-search artifacts, including JSON-LD, `llms.txt`, and Markdown output. Those artifacts are part of Strand's opinionated publishing core. Ghost's [documentation](https://ghost.org/docs/) describes its publishing and theme surface, and the [Content API](https://ghost.org/docs/content-api/) documents read access for delivery architecture. Ghost is therefore a credible headless option, but headless delivery does not turn Ghost into a Git-native CMS. ## Cost and responsibility Ghost(Pro) is a hosted subscription whose current plan boundaries and prices belong on [Ghost's pricing page](https://ghost.org/pricing/). Self-hosted Ghost uses the official open-source code, but the operator still pays for infrastructure, maintenance, backups, and operations. Strand's repository provides the software without a hosted subscription claim. That does not make a Strand publication free: hosting, domain, CI, design, monitoring, and editorial labor still cost something. Compare total responsibility, not just the license line. ## Choose Ghost if these conditions are true - Memberships or newsletters are central to the business. - Nontechnical editors need a hosted publication workspace. - You prefer Ghost's integrated product over assembling a Git pipeline. - A Content API or theme-based site is enough for delivery. ## Choose Strand if these conditions are true - Agents create or update content under a strict schema. - Pull-request-style review is more useful than a WYSIWYG editor. - MDX, Git history, and deterministic checks are requirements. - You want generated SEO and AI-search artifacts in the publishing core. ## Why this matters The important question is not “which CMS has more features?” It is “where should the publication's source of truth and review gate live?” Ghost is a strong publication product. Strand is a focused implementation for a Git-native, agent-first operating model. Read the [best Ghost CMS alternatives pillar](/blog/best-ghost-cms-alternatives), then see [Introducing Strand](/blog/introducing-strand) for the product context. ## Programmatic publishing scenarios Consider three common scenarios. In the first, an editorial team publishes a weekly newsletter and sells memberships. Ghost's integrated product is likely the shortest path because it directly serves that business model. Moving to a Git-native CMS would require replacing capabilities that are already central to the publication. In the second, a technical team creates a large set of sourced guides from a repository and wants every change reviewed before release. Strand's model is a closer fit because the content, metadata, and validation can live beside the code. The team must still write the policy and maintain the pipeline; automation does not remove that responsibility. In the third, a company needs one product catalog or knowledge model across a marketing site, application, and mobile client. A headless platform may be more appropriate than either integrated publication tool. Ghost can deliver content through its API, but the team should compare the data model and application requirements directly rather than assuming a blog API is a complete content platform. ## A migration decision checklist Before moving from Ghost, answer these questions in writing: - Which Ghost features are business-critical, and who replaces each one? - Which URLs, authors, tags, images, and embeds must survive? - Who will operate hosting, email, backups, and incident response? - Is the editorial team comfortable reviewing files or does it need a visual editor? - What deterministic checks must pass before a post can ship? - What is the rollback path if the first release has broken links? This checklist often produces a better answer than a feature count. “Stay on Ghost” can be the correct conclusion when it owns the requirements. “Move to Strand” can be the correct conclusion when Git review and agent-operated publishing are more valuable than a hosted editor. ## The comparison in day-to-day work Imagine an editor correcting a sourced paragraph. In Ghost, that person works in the publication interface and uses the product's publishing controls. In Strand, the correction is a file change: the diff shows exactly what changed, validation checks the contract, and the release process can require an approval. Neither path is automatically safer. The safety comes from whether the team follows the path and whether the system makes mistakes visible. Now imagine a developer changing the site template. Ghost's theme and API model can be useful when the publication wants a recognizable CMS boundary between content and presentation. Strand's repository model puts content and application work close together, which can reduce coordination for a technical team but increase the need for a disciplined deployment setup. For a content operation publishing many related pages, the difference compounds. A Git-native batch can make the complete set and its internal links inspectable before release. A Ghost operation may be faster for a single human-authored post. The right comparison is the unit of work your team repeats most often. ## How to run a fair proof of concept Use the same small corpus in both systems: a short post, a long guide, a comparison table, a post with an image, and a post with an external source. Have the same person perform the same tasks: draft, edit, preview, publish, correct, and restore an earlier version. Record friction instead of relying on impressions. Count the steps only if the count explains something; more important are the failure modes. Can a bad slug ship? Can a broken link be found before release? Can a reviewer tell what changed? Can a new operator understand the process from the documentation? Do not use an invented benchmark to declare a winner. A proof of concept is local evidence about your workflow. Publish the result as a decision record so the next team member knows why Ghost or Strand was selected. ## Editorial and operational fit The comparison also changes depending on who approves a post. A small engineering-led team may prefer a pull request because the diff, source links, and metadata are visible in one place. A newsroom with many nontechnical contributors may prefer Ghost's editor because the interface is the shortest route from assignment to publication. Consider corrections. If a published claim is wrong, Strand's Git history can make the correction traceable and the validator can protect the shape of the replacement. Ghost's publication workflow can make a human correction quick in the editor. The deciding factor is not which system has a theoretical advantage; it is which team can make and verify corrections under pressure. Consider media too. A programmatic blog may contain diagrams, code, screenshots, and external embeds. Test the rendering and asset workflow in both products. Do not treat a text-only proof as evidence that the whole publication can move cleanly. ## What each product asks the team to learn Ghost asks the team to learn its publication model, themes, integrations, and hosting boundary. That is a focused learning path for a publication, especially on Ghost(Pro). Strand asks the team to learn Git, MDX, schema validation, deployment, and agent operations. That is a better fit for developers who want content changes to behave like code changes. Neither learning path is free. Training, documentation, review time, and operational confidence belong in the migration decision. A technically elegant system can still fail if the people responsible for it do not want to run it. ## A fair conclusion Ghost and Strand are optimized for different centers of gravity. Ghost is a publication product whose strengths include an editor, memberships, newsletters, themes, and an API. Strand is a repository-first publishing system for teams that want content and validation close to code. Programmatic publishing does not erase that difference; it makes the review and operating model more important. Choose the system you can keep accurate after the launch announcement. That means sources remain attached, corrections are possible, URLs remain stable, and someone owns the next failed deployment. A narrow product that is operated well beats a broad product that no one understands. ## Operating a complete programmatic cluster A fair programmatic test should include more than one post. Publish a small related cluster and inspect whether the system preserves topic boundaries, internal links, source lists, and consistent metadata. This is where a Git-native workflow can be useful: the batch can be reviewed as a set before release, while a human publication workflow can use its own editorial calendar and approval process. Do not confuse speed of generation with quality of publication. The team still needs a research note, a source policy, a correction route, and a decision about when an article is complete. Ghost and Strand can both be operated with discipline; the tool does not excuse the absence of it. A useful operating review asks whether the team can explain one complete release from assignment to reader. Who researched it, which sources were checked, who approved it, what changed during review, and how would the team correct it tomorrow? If the answer is visible in the repository, a ticket, or the publication workflow, the process is teachable. If it lives only in one person's memory, the process is fragile. A final check is reversibility. Before choosing either product, make a test correction, make a test rollback, and document the exact steps. The workflow that is easiest to explain under normal conditions is usually the one the team can recover with when conditions are not normal. That is the standard to use when the headline question is “which wins?” The winner is the product whose workflow matches the team's actual skills, content risk, and business requirements—not the product with the longest feature list. In practice, the decision is often reversible at the beginning and expensive later. Run the proof while the corpus is small, preserve an export, and make the owner of the final decision explicit. That is more useful than waiting for certainty a feature page cannot provide. Write the decision down with the date and assumptions. Revisit it when the audience, team, or publication model changes. Keep the proof corpus and its results with that decision record so a later review compares the same facts rather than a new impression. ## FAQ ### Is Strand or Ghost better for programmatic publishing? Strand is the closer fit when content is generated and reviewed as Git changes under a schema. Ghost is better when automation must coexist with its integrated publication features. ### Can Ghost be used headlessly? Yes. Ghost's Content API documents read access to published content for separate front ends. ### Does Strand replace Ghost memberships? No. Strand is a Git-native publishing system, not a membership and newsletter product. ### Which costs less? Check Ghost(Pro)'s current pricing. Strand's software can be run without a software subscription, but hosting and labor remain costs. ## FAQ **Is Strand or Ghost better for programmatic publishing?** Strand is the closer fit when content is generated and reviewed as Git changes under a schema. Ghost is the better fit when programmatic content still needs Ghost's integrated editor, memberships, and publication features. **Can Ghost be used headlessly?** Yes. Ghost documents a Content API for reading published content, so teams can deliver Ghost content to a separate front end. **Does Strand replace Ghost memberships?** No. Strand is a Git-native publishing system, not a claim of feature parity with Ghost's membership and newsletter product. **Which costs less, Strand or Ghost?** Strand's open-source software can be run without a software subscription, but hosting and labor remain costs. Ghost(Pro) pricing is published by Ghost and should be checked for the current plan. ## Sources - [Strand CMS repository](https://github.com/BowTiedSwan/strand) — Strand CMS - [Ghost documentation](https://ghost.org/docs/) — Ghost - [Ghost Content API documentation](https://ghost.org/docs/content-api/) — Ghost - [Ghost(Pro) pricing](https://ghost.org/pricing/) — Ghost --- # What an Agent-First CMS Actually Cuts: The 85% Hook *By The Strand CMS Team · 2026-08-10 · 6 min read* Canonical: https://www.strandcms.com/blog/what-agent-first-cuts > **Summary:** The 85% in this agent-first CMS headline is a heuristic, not measured research. The real cut is the human-editor and database machinery; source checking, review, design, hosting, and maintenance still remain. “Agent-first CMS” describes an operating model, not a magic productivity statistic. The **85%** in this title is an editorial hook and heuristic, not measured research. It means that a repository-first team may not need much of the conventional editor-and-database machinery. It does not mean 85% less work, 85% lower cost, or 85% better output. The [Strand repository](https://github.com/BowTiedSwan/strand) describes the specific version used here: MDX content in Git, a validated schema, generated publishing artifacts, and agent tooling. The comparison is against a conventional publication CMS such as the one described in [Ghost's documentation](https://ghost.org/docs/), not against every CMS product. ## Agent-first CMS: what the 85% heuristic points at A conventional CMS often bundles a browser editor, content storage, roles, revisions, media handling, preview, publishing controls, and integrations. An agent-first stack can move several of those responsibilities into files, pull requests, scripts, and agent tools. That is a change in control plane. The content still needs a schema, an output renderer, a deploy target, and people who decide whether it should ship. The heuristic is useful only when it helps identify which machinery the team can intentionally replace. ## 1. The WYSIWYG editor is no longer the default An agent-first workflow can make Markdown or MDX in a repository the primary authoring surface. The benefit is inspectable text, normal diffs, and a format that automation can transform without driving a browser. The cost is that nontechnical editors may prefer a visual workspace. If a team needs inline layout controls, collaborative editing, or a marketer-friendly preview, removing the editor may create friction rather than remove it. “No WYSIWYG” is a trade-off, not a virtue in every organization. ## 2. The database stops being the source of truth Git can hold versioned content files and their history. In Strand's model, the repository is the source of truth for MDX posts, while validation protects the content contract before release. This can simplify backups, branching, rollback, and code review. It also means the team must understand Git, deployment, media storage, and merge conflicts. A database did not disappear; its responsibilities were redistributed to files and the surrounding toolchain. ## 3. Manual publishing orchestration can become a tool call Agent skills or an MCP server can create drafts, validate frontmatter, and prepare a publication batch. That reduces repetitive clicking and makes a release sequence explicit. Automation does not decide whether a source is credible or whether a claim is fair. The publishing tool should enforce deterministic checks, while editorial policy handles judgment. A system that skips those two layers is an automated slot machine, not an agent-first newsroom. ## What agent-first does not cut ### Review and source checking Google's [people-first content guidance](https://developers.google.com/search/docs/fundamentals/creating-helpful-content) is a useful quality floor: content should serve people rather than exist only to manipulate search. Agents can collect and transform evidence, but someone must define the evidence standard and handle uncertainty. ### Design and reader experience A clean MDX file does not design a useful site. Information architecture, typography, accessibility, navigation, mobile behavior, and media treatment still need ownership. ### Hosting and operations Git is not hosting. You still need a deploy target, domain, TLS, monitoring, backups, and an incident path. Open-source software can remove a license bill while leaving the operational bill intact. ### Corrections and accountability A validator can catch a missing field. It cannot know whether a competitor comparison is fair, whether a source supports a sentence, or whether a correction should be prominent. Those are editorial responsibilities. ### Distribution and demand An [llms.txt proposal](https://llmstxt.org/) can help describe a site's important pages, but a machine-readable index does not create demand or guarantee citations. Search, email, partnerships, and reader trust still require work. ## The hidden work an agent-first CMS exposes Replacing a CMS surface does not remove the decisions underneath it. A conventional product may hide those decisions behind defaults and screens. In a repository-first workflow they become explicit files, checks, and ownership boundaries. **Content modeling remains work.** Someone must decide which fields are required, how authors are represented, what qualifies as a source, and which states are allowed. A schema makes those rules executable, but it does not choose good rules for you. **Editorial planning remains work.** An agent can draft six related posts quickly, yet a cluster can still fail if every article says the same thing. Topic boundaries, internal links, canonical pages, and a complete release set need a human-owned plan. **Operations become more visible.** The team must know how a branch is created, how validation runs, how a failed check blocks release, and how a correction moves from an edited file to production. Visibility is a benefit only when the team documents and practices the path. **Readers remain the judge.** Agents may produce valid MDX that is boring, repetitive, or poorly matched to intent. The quality floor is still a useful answer, clear evidence, and a page that works for people first. ## What to keep in an agent-first stack Keep the parts that make publication reliable: - A strict schema with useful failure messages - Reviewable source files and history - Deterministic validation and link checks - A clear release gate for complete clusters - Human ownership of sources, corrections, and voice - A rendering layer that serves readers first This is the durable core behind Strand's product framing. It is also why a plain Markdown repository can be an honest alternative for a small technical site: the stack should fit the operator, not win a feature-count contest. ## Why this matters Agent-first design is valuable when it removes accidental complexity without removing accountability. Be precise about the cut: editor surface, database dependency, and manual orchestration may shrink; editorial judgment and operations do not. For the product context, see [Introducing Strand](/blog/introducing-strand). For the alternatives map, read [best Ghost CMS alternatives](/blog/best-ghost-cms-alternatives). ## The boundary between automation and judgment A validator answers structural questions: is the slug valid, is the author known, and does the frontmatter satisfy the contract? It cannot answer whether a claim is supported, whether a comparison is fair, or whether a reader will understand the opening. Those questions belong in editorial review. An agent-first CMS should therefore make review easier, not pretend review is obsolete. Store sources with the post, make changes diffable, and require a complete batch gate where that matches the publication's risk. The point is a faster path to a trustworthy release. ## When a conventional CMS is the better choice Keep a conventional CMS when the editorial surface is itself the product requirement. A visual editor can be the right answer for a team that does not want to learn Git. Ghost's [documentation](https://ghost.org/docs/) describes strengths around publication, themes, memberships, and APIs; removing those features is not an automatic improvement. Likewise, a database-backed system may be preferable when many users need concurrent structured editing, granular permissions, or application-driven content mutations. Agent-first is a fit criterion, not a moral ranking of architectures. ## FAQ ### What does an agent-first CMS remove? It can remove a WYSIWYG editor as the primary authoring surface, a content database as source of truth, and repetitive publishing orchestration. ### Does agent-first remove human review? No. Source checking, editorial judgment, corrections, and release gates remain necessary. ### Is 85% measured? No. It is a clearly labeled heuristic for replaced CMS machinery, not a measured cost or productivity statistic. ### What does Strand keep? A validated content schema and publishing core, with Git and agent tooling around the writing and release workflow. ## FAQ **What does an agent-first CMS remove?** It can remove a WYSIWYG editor as the primary writing surface, a content database as the source of truth, and some manual publishing orchestration. **Does agent-first mean no human review?** No. Source verification, editorial judgment, corrections, and release gates remain necessary. **Is the 85% figure measured?** No. It is an editorial heuristic for the portion of conventional CMS machinery a repository-first workflow may not need, not a measured statistic. **What does Strand keep?** Strand keeps a validated content schema and publishing core while using Git and agent skills for the writing and release workflow. ## Sources - [Strand CMS repository](https://github.com/BowTiedSwan/strand) — Strand CMS - [Ghost documentation](https://ghost.org/docs/) — Ghost - [Ghost Content API documentation](https://ghost.org/docs/content-api/) — Ghost - [Creating helpful, reliable, people-first content](https://developers.google.com/search/docs/fundamentals/creating-helpful-content) — Google Search Central - [The llms.txt proposal](https://llmstxt.org/) — llmstxt.org --- # Agent Publishing Safety: Direct Pushes with Guardrails *By The Strand CMS Team · 2026-08-10 · 6 min read* Canonical: https://www.strandcms.com/blog/agent-publishing-safety > **Summary:** Agent publishing safety comes from independent controls around the model: scoped permissions, verified sources, schema checks, branch policy, and a complete-batch gate. **Agent publishing safety** is not a warning label added after automation is built. It is the design of the path from source to production. MCP provides a structured way to invoke tools, Claude Code can operate in a repository, and GitHub provides repository rules. None of those facts makes an article sourced or a release complete by itself. This article supports the [best MCP server publishing](/blog/best-mcp-for-publishing) pillar. Read [publish a blog with Claude Code](/blog/publish-blog-claude-code) for the operator workflow and [Introducing Strand](/blog/introducing-strand) for the product context. ## Agent publishing safety: the control stack Use independent controls in layers: 1. **Identity:** know which agent, profile, or credential is acting. 2. **Scope:** limit files, tools, repositories, and environments. 3. **Evidence:** require real sources for factual claims. 4. **Validation:** reject malformed frontmatter and broken structure. 5. **Review:** inspect the diff and fairness of comparisons. 6. **Release:** publish only a complete, validated batch. 7. **Recovery:** preserve a reversible source history. The layers are deliberately redundant. If a model misunderstands a task, a schema check can still reject the result. If a tool is misconfigured, branch policy can still stop the push. ## MCP is not authorization The MCP tools documentation explains how servers expose discoverable functions with input schemas. That makes tool use clearer, but “the model can see a tool” is not the same as “the model is authorized to perform every consequence of that tool.” Keep read, write, validate, and publish permissions distinct where possible. A draft tool can write a branch. A publish tool should require a batch identifier and verify the final state. A destructive operation should be absent if the workflow does not need it. ## Require a source-to-claim path The most serious editorial failure is a fluent sentence that no source supports. Build the research note before the draft. Map every number, date, price, named feature, and quote to a URL that was actually retrieved. Product comparisons need extra care. If a vendor documents an API, say API. Do not turn that into “official MCP server” unless the vendor or protocol repository documents it. If a competitor wins on memberships, ecosystem breadth, or structured modeling, say so. ## Put deterministic checks outside the model A validator should inspect the file itself. The check can enforce required frontmatter, dates, slugs, author references, and schema types. Batch QA can enforce grounding, keyword placement, word bands, internal links, sources, and complete-cluster links. These checks are valuable precisely because they do not negotiate. A model can explain why a missing source is probably fine; the checker should still fail. ## Direct push does not mean blind push A direct push can be a safe release mechanism if the command verifies the branch and the batch before moving `main`. The required conditions should include: - the batch branch is based on current `origin/main` - every planned article exists - each article has the target status and date - schema validation passes - batch QA passes - only the intended content files are staged - the operation is a fast-forward If any condition fails, stop. Never use a force push to turn an unsafe state into a green-looking history. ## Use repository policy as a second line GitHub documents rulesets for repository-level enforcement. Rules can require status checks, restrict updates, or control who can push. The exact configuration depends on the repository, but the principle is stable: a publication agent should not be the only system deciding whether its own write is acceptable. The repository should retain a readable trail: research note, one commit per article, validation output in the run log, and the final batch publication. ## What to do when a gate fails **Bad source:** replace or remove the claim. **Schema error:** fix the frontmatter or stop if the schema contract conflicts with the editorial plan. **Broken link:** correct the slug or remove the link. **Incomplete batch:** finish the missing article; do not publish the partial set. **Branch drift:** update the branch safely from current `main`, preserving the work. **Permission error:** reduce scope or request the right access; do not work around it. ## The approval boundary The most important boundary is the moment a repository change becomes a published page. Treat that as an approval boundary with a named condition, not as the natural final step of a conversation. The condition can be a human review, a required check, or both, depending on risk. For a factual article or comparison, require evidence review. For a routine formatting change, deterministic checks may carry more of the burden. The policy should say which is which instead of pretending every edit has identical risk. ## Idempotency and retry behavior Agents retry. Networks fail, commands time out, and a client may not know whether a write completed. A safe publishing tool should make retries predictable: identify the target by slug or batch, detect an existing draft, and refuse an ambiguous overwrite. Return a durable result that the agent can inspect. The same principle applies to publication. The command should verify the current branch and target state before pushing. If the state changed underneath it, stop and report drift rather than guessing which history is safe. ## Observability without secrets Keep enough evidence to reconstruct a run: phase, branch, changed files, check results, and publication target. Do not put credentials, tokens, or private source material into public article text or commit messages. Logs should support accountability without expanding the exposure of the tool connection. A small run report is often enough. It can list each slug, word count, source count, validation state, and any unresolved claim. “No unresolved claim” should mean the claim was checked or removed, not that the model failed to mention uncertainty. ## Threats worth modeling The useful threat model is ordinary. A prompt may ask the agent to publish an unrelated file. A source may contain text that tries to redirect the workflow. A retry may create a duplicate draft. A stale branch may overwrite newer work. A vendor page may change after the article was written. A credential may be broader than the job requires. Treat external pages as evidence, not instructions. Treat repository state as something to inspect, not something to infer. Treat every production side effect as a separate authorization decision. ## Make the release report boring After a successful run, report the batch branch, article slugs, commit sequence, validation state, QA state, source counts, and unresolved claims. A failure report should name the command, the failing check, and the files affected. Do not hide a blocker behind a general “the agent encountered an issue.” Boring reports support fast human verification. They also make scheduled runs comparable over time without turning the publication into an opaque autonomous system. ## Separate risk by action Reading a source, writing a draft, changing metadata, and pushing production do not have the same risk. Give each action an appropriate control rather than applying one broad permission to the whole run. This makes the policy easier to explain and the incident smaller when something goes wrong. ## Why this matters Safety is what makes direct publication boring. The ideal release is an observable fast-forward whose inputs were sourced, whose shape was validated, and whose completeness was checked before the agent was allowed to move the branch. ## FAQ **What is agent publishing safety?** It is the set of controls that keeps an AI-assisted publishing workflow from making unsupported, malformed, unauthorized, or partial releases. **Are direct pushes unsafe by definition?** No. A direct push can be controlled by branch ancestry, required checks, scoped credentials, and a publication command that verifies the complete batch. **What should an agent never bypass?** It should never bypass failed validation, source verification, branch policy, or the requirement that every planned article exists before publication. ## Sources - [Model Context Protocol Specification](https://modelcontextprotocol.io/specification/2025-06-18) — Model Context Protocol - [Tools - Model Context Protocol](https://modelcontextprotocol.io/docs/concepts/tools) — Model Context Protocol - [Claude Code MCP](https://docs.anthropic.com/en/docs/claude-code/mcp) — Anthropic - [About rulesets](https://docs.github.com/en/repositories/configuring-branches-and-merges-in-your-repository/managing-rulesets/about-rulesets) — GitHub - [Strand CMS](https://github.com/BowTiedSwan/strand) — BowTiedSwan --- # Building an AI Content Agent That Doesn't Produce Slop *By The Strand CMS Team · 2026-08-10 · 9 min read* Canonical: https://www.strandcms.com/blog/ai-content-agent-no-slop > **Summary:** An AI content agent avoids slop when research, drafting, validation, and publication are separate stages with explicit evidence and a rejection path. An **AI content agent** is not a prompt with a publish button. It is a workflow that gives a model a bounded job, a set of evidence, and a way to fail safely. Anthropic's guidance on effective agents emphasizes simple, composable patterns before adding more autonomy. Google likewise frames helpful content around people, not mass-produced filler. This support article belongs to the [best MCP server publishing](/blog/best-mcp-for-publishing) pillar. See [agent publishing safety](/blog/agent-publishing-safety) for the release controls and [Introducing Strand](/blog/introducing-strand) for the Git-native product context. ## AI content agent: the quality model Quality is a chain, not a personality trait: `brief → sources → draft → checks → editorial review → publish` If one link is missing, the model can produce fluent but weak output. A source list without claim checking creates citation theater. A validator without editorial review catches shape but not meaning. A human review without deterministic checks wastes attention on preventable errors. ## 1. Give the agent a narrow assignment A useful brief states the audience, search intent, target keyword, content type, word band, internal links, publication date, and forbidden claims. It also says which files are in scope. Narrow scope makes the diff legible and gives the agent fewer ways to “help” by changing unrelated files. Do not ask for “a definitive article on everything.” Ask for one answer that a reader can use. If the subject needs several answers, make a cluster and give each article a separate job. ## 2. Make evidence a required input Extract checkable claims before drafting: numbers, dates, prices, named features, quotes, and legal or technical statements. For each claim, record a source that actually supports it. Prefer first-party documentation for product capabilities and primary research for study findings. If a claim cannot be verified, the options are limited: remove it, soften it to analysis, or attribute it explicitly as a report that still needs review. “The model probably knows this” is not a source. ## 3. Use narrow tools instead of broad authority MCP tools can expose structured operations with input schemas. That is a better fit for editorial systems than a single unrestricted “manage content” command. Give the agent tools such as: - create a draft with required frontmatter - read a source or research note - validate one post - run batch QA - show the diff - request publication of a complete batch The tool boundary should make the safe path easier than the unsafe path. A tool that accepts arbitrary database patches may be flexible, but it is harder to audit than a typed editorial operation. ## 4. Separate content checks from taste Deterministic checks should reject what a machine can decide: invalid YAML, missing fields, wrong slugs, bad dates, missing links, absent grounding, and word counts outside the band. Editorial review should focus on what requires judgment: whether the answer is useful, whether the sources support the claims, whether the comparison is fair, and whether the language is repetitive. Do not use a model's self-critique as a replacement for either category. A second generated paragraph can be helpful; it is not evidence that the first paragraph is true. ## 5. Detect common forms of slop Look for: - generic openings that could fit any industry - repeated conclusions with no new information - feature lists that never explain the reader's outcome - confident claims with no source - fake precision and round numbers - competitor summaries that omit their strengths - keyword repetition that makes the page harder to read - headings that promise an answer the section never gives The cure is usually subtraction. Cut the sentence, narrow the claim, add the missing source, or replace the abstraction with a concrete example. ## 6. Keep the editor in the loop where it matters The model should not decide whether its own unsupported claim is acceptable. A reviewer should be able to see the sources, the exact diff, the validation output, and the publication target. If a human is not available for every low-risk edit, make the deterministic envelope stronger and reserve human attention for claims, comparisons, and release decisions. This is not anti-agent. It is what makes agentic work operationally trustworthy. ## 7. Build a rejection path Every stage needs a clear failure result. A source fetch can fail. A schema check can fail. A link can point to a missing slug. A branch can be behind `main`. The system should stop and report the blocker instead of silently producing a partial cluster. A rejection path is also a design signal. If the workflow cannot explain how to undo a write, it is not ready for autonomous publication. ## 8. Use a complete-cluster gate A pillar without its supports is not a finished editorial unit. The pillar should link every support, while every support links back to the pillar and at least one money page. Check that every planned file exists, uses the same target date, and has the final status before publishing. Batching also makes review more coherent. The editor can assess whether the articles cover distinct intents rather than repeating the same answer six times. ## A practical agent checklist Before the first write: - [ ] branch is based on current `main` - [ ] brief names the exact files in scope - [ ] sources are real and retrieved - [ ] unsupported claims are marked for removal Before publication: - [ ] every article starts with `` - [ ] frontmatter validates - [ ] keyword placement remains natural - [ ] sources are visible - [ ] pillar and money links exist - [ ] batch QA passes - [ ] diff has been reviewed ## Reusable prompt contract A good agent brief can be reused across a week of posts without becoming a blank check. State the reader, the article's one question, the allowed sources, the required links, the frontmatter contract, and the commands that must pass. State what the agent must not do: no invented numbers, no unsourced competitor claims, no unrelated file changes, and no publish call before the batch gate. The brief should also require an answer-first opening. A reader should learn the practical recommendation before meeting the process detail. That structure helps the editor spot whether the draft actually answers the query or merely circles it with fluent context. ## Review the output at three levels Review the sentence level for unsupported facts and awkward certainty. Review the section level for one idea per heading and useful examples. Review the cluster level for distinct search intent and a coherent link graph. A draft can pass one level and fail another. At sentence level, ask what evidence supports the claim. At section level, ask whether the reader can act on the advice. At cluster level, ask whether the support articles add coverage or simply repeat the pillar. These questions are faster than trying to judge “quality” as one large feeling. ## Add correction mechanics A responsible workflow needs a correction path before the first release. Keep the source commit, source note, and update date visible. When a source changes, update the claim, its citation, and the article's mutable metadata together. Do not quietly rewrite a disputed sentence without recording why. Corrections are also useful feedback for the agent. If a source was overread, add that failure to the editorial procedure. If a tool allowed a dangerous state, narrow its schema or permissions. The goal is not to promise zero mistakes; it is to make mistakes detectable and reversible. ## Build a claim ledger For a serious article, keep a small claim ledger beside the draft. Each row can contain the claim, claim type, source URL, exact supporting passage or section, and editorial treatment. Claim types include product capability, technical behavior, date, price, statistic, and opinion. This turns “fact checking” into a visible task instead of a final mood. A ledger also helps with updates. When a vendor changes a feature, the editor can find the affected sentences instead of rereading every paragraph from scratch. Mutable claims deserve a date in the prose when the date changes what the reader should conclude. ## Prefer useful specificity Avoid both empty generalities and invented precision. “Agents need guardrails” is a starting point, not a useful conclusion. Explain which guardrail catches which failure: a schema catches missing fields, a source note catches unsupported claims, a branch check catches stale history, and a batch gate catches partial clusters. Specificity does not require a made-up statistic. It can come from a reproducible procedure, a concrete file path, a real command, or a clearly labeled tradeoff. ## Make the output easy to edit Short sections, meaningful headings, tables with consistent columns, and visible source links help the editor work quickly. They also make it easier to compare the draft with the brief. If the agent produces a long paragraph that mixes evidence, recommendation, and caveat, split it into separate units. A good article should still be useful if the reader ignores the product mention. Product relevance comes from the workflow problem the product actually solves, not from repeating the brand name. ## Use a stopping budget Give the agent permission to stop after a bounded number of retries or when evidence is missing. Endless rewriting often makes a weak claim more elaborate without making it more defensible. A clear stop lets an editor decide whether to narrow the brief, find a better source, or remove the topic. ## Treat quality as a product requirement A publication should define quality before automation begins. That definition can include answer-first structure, visible sources, a fair competitor section, a useful internal link graph, and a correction path. The agent then has a target that is more concrete than “sound expert.” The standard should be applied to machine-assisted copy and human copy alike. Otherwise the workflow teaches the agent that polished language can excuse weak evidence, which is precisely the habit the guardrails are meant to prevent. ## Give reviewers a compact surface Reviewers should not have to search an entire run to understand a proposed release. Put the source note, changed slugs, validation result, QA result, and unresolved claim list in one report. Link the report to the commits and article files. This lets a reviewer spend time on evidence and meaning instead of reconstructing mechanics. A compact surface also makes disagreement productive. A reviewer can point to the exact source, sentence, or gate rather than saying that the article “feels AI-generated.” The agent can then revise a claim or remove a section with a reason attached. That is a better loop than asking for another unstructured rewrite. It leaves a useful editorial record for the next update. The record is part of the quality system. It lets the next editor start from evidence and keeps the revision explainable. That makes the next pass faster. It also protects consistency. It keeps editorial decisions visible. ## Why this matters An AI content agent is useful when it makes the evidence and the decision points clearer, not when it hides them behind polished prose. Separating stages gives the team a practical way to scale output without scaling unsupported certainty. ## FAQ **What is an AI content agent?** It is a software workflow in which a model uses tools and instructions to research, draft, check, or move content through a publication process. **How do you stop AI content from becoming slop?** Require real sources, narrow the assignment, validate structure deterministically, review the diff, and block publication when claims or links fail. **Should an AI agent publish without review?** Only when the surrounding system has independently enforced the required checks and the risk is acceptable; model confidence is not a quality gate. ## Sources - [Building effective agents](https://www.anthropic.com/research/building-effective-agents) — Anthropic - [Creating helpful, reliable, people-first content](https://developers.google.com/search/docs/fundamentals/creating-helpful-content) — Google Search Central - [AI features and your website](https://developers.google.com/search/docs/fundamentals/ai-overviews) — Google Search Central - [Tools - Model Context Protocol](https://modelcontextprotocol.io/docs/concepts/tools) — Model Context Protocol - [Strand CMS](https://github.com/BowTiedSwan/strand) — BowTiedSwan --- # Publishing MCP Tools: create_draft to publish_post *By The Strand CMS Team · 2026-08-10 · 7 min read* Canonical: https://www.strandcms.com/blog/anatomy-publishing-mcp > **Summary:** Publishing MCP tools work best as a narrow lifecycle: create a draft, validate the exact file, run complete-batch QA, and publish only after independent checks pass. The useful shape of **publishing MCP tools** is a lifecycle, not a single super-tool. The Model Context Protocol documentation describes discoverable tools with structured inputs and results. A content system can use that contract to expose small editorial operations: create the file, validate the file, check the cluster, then publish the approved state. This is a support article for [Best MCP Server Publishing Tools in 2026](/blog/best-mcp-for-publishing). It also connects to [Claude Code's full blog pipeline](/blog/publish-blog-claude-code), [AI content agents without slop](/blog/ai-content-agent-no-slop), and [Introducing Strand](/blog/introducing-strand). ## Publishing MCP tools: the lifecycle ```text brief → create_draft → validate_post → batch QA → review → publish_post ``` Each arrow matters. The next tool should consume the actual result of the previous step rather than a model's memory of what happened. ## 1. create_draft `create_draft` should accept a slug, frontmatter, and body—or a similarly explicit typed payload. It should write only to the permitted working area and return a durable identifier such as the path or slug. A good draft tool rejects malformed input early. It can check that the slug matches the filename, that an author exists, and that required fields are present. It should not silently convert a missing source into a complete article. The tool should also be idempotent or explicit about overwrites. “Create” should not unexpectedly destroy an existing article because the model retried a request. ## 2. validate_post `validate_post` checks the exact on-disk file against the content schema. In a Strand workflow that includes title and description limits, ISO dates, author references, content metadata, FAQ entries, sources, and the first `` component. The key word is exact. Validate the file that will be committed, not an in-memory draft that might differ from the repository. Return actionable errors with a field or line where possible. A successful schema check is necessary, not sufficient. It does not verify that a source supports a claim or that a comparison is fair. ## 3. Batch QA A post can validate alone and still fail as part of a cluster. Batch QA checks the link graph, keyword placement, word band, money-page link, sources, and the requirement that the pillar links every support. It also catches a missing planned file. This is why `publish_post` should not be called per article when the editorial unit is a cluster. The batch is the thing that needs to be complete. ## 4. Review Review is where the editor reads for meaning. Compare each factual sentence with its source. Check that the article answers its primary question in the opening. Look for repeated paragraphs, generic AI language, and competitor descriptions that hide a genuine strength. A useful review record includes the source note, diff, validation output, and the intended publication date. That is enough context to make a release decision without reconstructing the whole prompt history. ## 5. publish_post A publish operation should be the narrowest and most protected tool. Its inputs should identify the intended post or batch, and the server or wrapper should re-check status, branch ancestry, required files, and validation before the side effect. The MCP specification defines the interaction model; it does not prescribe your publication policy. Your wrapper must supply that policy. In a Git-native system, the final operation can be a fast-forward of a validated branch to `main`. ## Tool schemas should express the policy Avoid an input like: ```json {"action":"do_anything","payload":"..."} ``` Prefer explicit inputs such as: ```json {"slug":"example-post","status":"draft"} ``` or: ```json {"batch":3,"expectedSlugs":["pillar","support-a"]} ``` Schemas are not a full security model, but they make invalid states visible and help clients choose the correct tool. ## Permissions and failure states Separate credentials for draft and publish operations reduce blast radius. A read-only research client should not be able to push. A draft writer should not be able to bypass batch QA. GitHub rulesets can provide repository-level enforcement around updates and required checks. Design failures as normal results: source unavailable, invalid frontmatter, missing article, stale branch, failed QA, or denied permission. The agent should report the exact blocker and stop. ## Strand's concrete example Strand's repository is the source for its documented MCP tool names and Git-native workflow. The useful pattern is not the brand-specific names; it is the sequence. A structured draft is created, the exact post is validated, the batch is checked, and publication is a separate action. That sequence keeps the source visible and gives the editor several points at which to catch unsupported claims or accidental scope expansion. ## Designing the result object A tool result should help the next step decide what to do. For a draft, return the path, slug, and whether an existing file was changed. For validation, return pass or fail plus field-level errors. For batch QA, return every article and every failed check. For publication, return the target branch and commit or a clear failure. Avoid a result that says only “success.” The agent needs enough detail to avoid repeating a completed write or claiming that a different file was validated. A useful result is part of the audit trail. ## Content state is not tool state A tool can report that it wrote a file while the article remains a draft. A validator can report that the schema passed while the sources are weak. A publication command can report a push while the site deployment is still processing. Keep these states distinct in both the interface and the run report. This distinction prevents overclaiming. The editor can say “the batch branch passed validation and was pushed” without saying “every reader has seen the updated page” unless deployment has also been verified. ## A minimal lifecycle example Imagine a support article about Claude Code. The agent first receives a brief and source list. `create_draft` writes the MDX. `validate_post` checks the frontmatter and component order. Batch QA verifies the support links back to its pillar and includes a money-page link. The reviewer checks the Anthropic documentation against the claims. Only then does the batch-level publish operation run. If the title is too long, the schema fails. If an API claim lacks a first-party URL, editorial review fails. If the pillar does not link the support, QA fails. Each failure is better than a polished but incomplete release. ## What not to expose Do not expose arbitrary shell execution, unrestricted database writes, credential retrieval, or a tool that combines research, drafting, and publication into one opaque action unless the surrounding environment has a compelling reason and strong controls. Most editorial workflows need less power than that. Least privilege also improves usability. When a tool has one job, its description can be precise, its input schema can be small, and its error can point to a specific correction. ## Version the contract MCP specifications and vendor clients evolve. Keep the server's supported protocol version, tool descriptions, and input schemas visible to the project. When a tool changes behavior, test an existing article fixture before connecting it to the production workflow. Versioning is not overkill for a publication. It lets the editor distinguish a content error from an integration change and gives the team a clear place to document compatibility. ## Test with fixtures Use a valid article, a missing-field article, a bad-link article, and an intentionally incomplete batch as fixtures. The expected result should be explicit for each one. A server that passes the valid fixture but also publishes the incomplete batch is not safe enough, regardless of how convenient the happy path feels. Fixtures also make refactoring less risky. When the content schema or tool wrapper changes, rerun the same examples and inspect the result rather than relying on a manual click-through. ## Why this matters A publishing MCP is an interface between an agent and a release system. Good interfaces reduce ambiguity. They make the safe path discoverable, make bad inputs rejectable, and make the final side effect rare enough to inspect. The best implementation is not the one with the cleverest tool name. It is the one that still refuses to publish when the article is incomplete. ## FAQ **What are publishing MCP tools?** They are MCP-exposed functions for bounded editorial actions, such as creating a draft, validating a post, or requesting a controlled publication. **Why separate create_draft and publish_post?** Separation keeps drafting reversible and lets validation, review, and permission checks happen before a production side effect. **Does validate_post prove an article is good?** It proves the checks it implements passed. Editorial judgment, source quality, fairness, and usefulness still require review. ## Sources - [Tools - Model Context Protocol](https://modelcontextprotocol.io/docs/concepts/tools) — Model Context Protocol - [Model Context Protocol Specification](https://modelcontextprotocol.io/specification/2025-06-18) — Model Context Protocol - [Claude Code MCP](https://docs.anthropic.com/en/docs/claude-code/mcp) — Anthropic - [Strand CMS](https://github.com/BowTiedSwan/strand) — BowTiedSwan - [About rulesets](https://docs.github.com/en/repositories/configuring-branches-and-merges-in-your-repository/managing-rulesets/about-rulesets) — GitHub --- # Best MCP Server Publishing Tools in 2026 — 6 Compared *By The Strand CMS Team · 2026-08-10 · 12 min read* Canonical: https://www.strandcms.com/blog/best-mcp-for-publishing > **Summary:** MCP server publishing is a protocol and workflow choice, not a magic ranking feature. Compare six practical paths by tool boundaries, source control, validation, and publication safety. The best **MCP server publishing** setup is the one that makes the next safe action obvious. MCP gives a client a standard way to discover tools, send structured inputs, and receive results. It does not decide whether a paragraph is true, whether a branch is protected, or whether a live push should happen. This is the pillar for [publishing a blog with Claude Code](/blog/publish-blog-claude-code), [AI content agents without slop](/blog/ai-content-agent-no-slop), [agent publishing safety](/blog/agent-publishing-safety), [the Hermes editor profile](/blog/hermes-editor-profile), and [the anatomy of a publishing MCP](/blog/anatomy-publishing-mcp). For the product itself, see [Introducing Strand](/blog/introducing-strand). ## MCP server publishing: quick verdict | # | Tool or approach | Best for | Source of truth | Main strength | Main limitation | |---:|---|---|---|---|---| | 1 | Strand MCP | Git-native editorial teams | MDX in Git | Draft, validate, and batch publication fit together | Narrower ecosystem than a general CMS | | 2 | Ghost API plus MCP bridge | Membership publications | Ghost database/API | Strong publishing and membership model | The bridge and permissions are your responsibility | | 3 | WordPress REST API plus MCP bridge | Existing WordPress sites | WordPress database | Huge ecosystem and mature REST surface | Plugin and schema drift need discipline | | 4 | Sanity API plus MCP bridge | Structured content teams | Hosted content lake | Flexible modeling and custom frontends | More infrastructure decisions around publishing | | 5 | Strapi API plus MCP bridge | API-first self-hosted teams | Application database | Admin UI plus extensible APIs | Validation and release policy sit outside MCP | | 6 | Plain Markdown plus a small server | Small technical teams | Repository files | Fewest moving parts | You must build editorial ergonomics | ## What MCP actually standardizes The MCP documentation describes tools as functions exposed by a server for a model to discover and invoke. Tool definitions include names, descriptions, and input schemas. A client can ask what is available, then call one tool with structured arguments and inspect the returned result. That is useful for publishing because the verbs can be narrow. `create_draft` is easier to audit than “manage the website.” `validate_post` gives the client a deterministic checkpoint. `publish_post` should still be the last step, not the default side effect of drafting. The protocol is a transport and interaction contract. It is not a content policy, a fact-checker, or a branch-protection system. ## The six approaches in detail ### Strand MCP Strand is the strongest fit when the team wants publication to remain close to the source. Its repository describes a Git-native CMS in which content is MDX and the publishing surface is produced from that source. Its MCP tools are designed around editorial operations rather than a generic database mutation. The useful distinction is not that an agent can call a tool. Many systems can expose an API. The distinction is that the call can sit inside a reviewable file workflow: write a draft, inspect the diff, validate the schema, run batch QA, then publish a complete cluster. Choose Strand if your priority is a clean repository history, explicit frontmatter, and deterministic gates. Choose another system if you need a mature membership business, a large plugin ecosystem, or a hosted editorial dashboard as the center of the operation. ### Ghost with an MCP bridge Ghost's Admin API is a documented administrative surface for managing content and other publication operations. Ghost is a genuine winner when the publication is also a newsletter or membership business. That product shape is a meaningful advantage, not a footnote. An MCP bridge can translate model requests into Ghost API calls, but the bridge becomes a security boundary. It needs scoped credentials, careful handling of content, idempotency, and a clear distinction between draft and published states. The API does not turn a model into an editor who has verified every claim. Choose Ghost if memberships, newsletters, and a polished publisher workflow matter most. Do not choose it solely because an MCP wrapper makes an API callable. ### WordPress with an MCP bridge WordPress's REST API is the practical route for organizations already operating WordPress. It has a broad content model and an enormous ecosystem. If a team has existing plugins, editorial roles, and distribution workflows, the cost of switching can outweigh the elegance of a new agent-native stack. The tradeoff is surface area. A bridge must know which post types, taxonomies, custom fields, and status transitions are allowed. It also has to survive plugin changes. WordPress wins on familiarity and integration breadth; it loses when a team wants a tiny, opinionated publishing surface. Choose WordPress if the site already depends on its ecosystem. Add validation and review outside the model call. ### Sanity with an MCP bridge Sanity is a strong option when structured content and custom frontends are the product requirement. Its documented content model and API-oriented approach give teams flexibility to represent more than articles: product data, reusable blocks, and editorial entities can share a system. That flexibility changes the agent problem. The agent needs to understand the schema, references, validation rules, and preview behavior. The MCP layer should expose domain operations rather than raw mutations wherever possible. Choose Sanity if one structured content model must serve several products. Choose a repository-first system if transparent diffs are more important than centralized structured content. ### Strapi with an MCP bridge Strapi documents REST APIs and a self-hosted CMS model. It is a reasonable fit for teams that want an admin panel, API-first delivery, and control over deployment. That combination is a real strength for application teams. The agent workflow needs an explicit release layer. An MCP server can create or update content, but it should not silently promote an entry to production. Use role-aware credentials, schema validation, previews, and a separate publication command. Strapi wins when the CMS is part of a broader application architecture. ### Plain Markdown in a repository Plain Markdown is the honest option for a small technical team that does not need a CMS UI. A thin MCP server can expose “write file,” “run checks,” and “open a review” operations. That can be enough for a modest publication. The downside is everything you did not build: previews, asset handling, author management, scheduling, link checks, and structured metadata. Markdown is not automatically better. It is better when the team understands the missing ergonomics and chooses the lower operational surface deliberately. ## The factors that matter more than the MCP label | Factor | Question to ask | |---|---| | Tool scope | Does each tool do one understandable thing? | | Input schema | Can invalid slugs, dates, and statuses be rejected before a write? | | Permission | Can drafting and publishing use different credentials? | | Preview | Can a human inspect the exact rendered result? | | Validation | Does publication run deterministic checks outside the model? | | Rollback | Can the team revert the source without reconstructing state? | | Audit trail | Is there a durable record of who changed what and when? | MCP is most useful when it shrinks ambiguity. A server that exposes every underlying database operation may be technically powerful but editorially unsafe. A smaller server with clear verbs is easier for both a model and a reviewer to use correctly. ## Choose the competitor if - Choose Ghost if subscriptions, newsletters, and memberships are central. - Choose WordPress if your existing plugin and agency ecosystem is a major asset. - Choose Sanity if one structured content model must serve multiple custom applications. - Choose Strapi if you want a self-hosted API-first CMS with an admin surface. - Choose plain Markdown if the publication is small and you want to own every layer. ## Choose Strand if - Your source of truth should be MDX in Git. - You want the pillar and support articles reviewed as one batch. - You want schema validation and content QA before a direct push. - Your agents should operate through narrow editorial tools. - You value a clean Markdown output alongside the HTML page. ## A simple scoring method A comparison is easier to trust when the criteria are visible. Score each approach on five questions: can the agent discover the operation, can the server reject malformed input, can a reviewer inspect the source change, can publication be separated from drafting, and can the team recover from a bad release? Do not pretend these are laboratory measurements. They are decision criteria. A tool that scores well on discovery but poorly on recovery is not a safe default. A system that scores well on recovery but requires extensive custom integration may still be the right fit for a technical team. The purpose of the table is to expose the trade, not to turn editorial architecture into a fake benchmark. ## What a team should document Before connecting an agent, write a small operating contract: - accepted content types and required fields - allowed repositories and branches - which tools are read-only - who or what can approve publication - how factual claims are sourced - what happens on a failed check - how a correction is made after release This contract prevents the most common category error in agent publishing: treating a tool invocation as the same thing as an editorial decision. The invocation is just an event. The policy determines whether the event is allowed to change the publication. ## API access versus publishing ownership A CMS API can be valuable without being an MCP server. Conversely, an MCP wrapper can make an API callable without improving the underlying editorial process. Ask where the source of truth lives, who owns the schema, and how a reviewer sees the exact output. Ghost's API is valuable because Ghost itself is a strong publication and membership product. WordPress's API is valuable because its ecosystem is broad. Sanity and Strapi are valuable when structured or application-oriented content is central. These strengths remain true even when an agent is not involved. Strand's advantage is narrower: it keeps the source, validation, and Git history close together for teams that want that model. Narrow advantages are often more useful than broad claims. ## How to run a fair trial Choose one representative article rather than a toy example. Give every candidate the same brief, sources, required fields, internal links, and review standard. Record setup work separately from writing quality. Test the failure path by supplying a missing field or unsupported claim and observe whether the system blocks the release. Do not publish the trial article simply because the prose looks polished. Read the sources, inspect the generated page, and check the rollback path. A short evaluation that exposes the operational tradeoffs is more useful than an invented scorecard. ## Testing the end-to-end path Do not evaluate a publishing tool only on its happy path. Run a small test matrix. First, submit a valid draft and confirm that the returned path is the one you expected. Next, remove a required field and confirm validation fails. Then change the slug and confirm the link checker catches the old URL. Finally, make the branch stale and confirm publication refuses to guess. The point is not to create theatrical failure. It is to see whether the system fails close to the cause. A useful tool tells the editor which input was rejected and what state remains on disk. A poor tool leaves a half-change and a vague success message. ## The hidden cost of a bridge An MCP bridge is software that needs its own maintenance. It may need to translate Markdown to a CMS's rich-text model, map status values, preserve authors and tags, handle retries, and keep credentials safe. Those costs do not make a bridge wrong. They make them part of the comparison. A bridge is easier to justify when the underlying CMS already holds valuable content and workflows. Starting a new publication gives you more freedom to choose a simpler source of truth. The honest question is not “can this connect?” but “what must we keep synchronized after six months?” ## Where each option wins Ghost wins when a publication's business model is memberships and newsletters. WordPress wins when the existing ecosystem is the constraint. Sanity wins when structured content must flow to several experiences. Strapi wins when a self-hosted API and admin interface fit the application. Plain Markdown wins when the team is small and comfortable owning the missing layers. Strand wins when the team wants the article source to remain a first-class review artifact. That is a specific operational advantage, not a claim that it replaces every CMS category. ## A final buyer question Ask which failure you are most willing to own. A hosted CMS may reduce publishing maintenance while a repository workflow may reduce source drift. A bridge may preserve existing investment while adding integration upkeep. The right answer is the one whose failure mode the team can actually operate. The practical recommendation is to test the failure path before you compare feature lists. If the system cannot explain what happens when validation fails, it is not ready to own a production publication. Also check whether the team can export or recover the source without depending on the same agent that made the change. That test usually reveals more than another feature checkbox. ## Why this matters MCP makes publishing callable. It does not make publishing correct. The durable advantage comes from pairing a standard tool interface with a source of truth, a fact-checking habit, and gates that the model cannot talk its way around. For a small team, that often means the best server is not the one with the most tools. It is the one whose next action can be explained in one sentence and verified in one command. ## FAQ **What is an MCP server for publishing?** It exposes publishing operations as structured tools that an MCP client can discover and call, such as creating a draft, validating content, or requesting publication. **Does MCP make an AI publisher safe by itself?** No. Safety still depends on permissions, validation, review, branch controls, and a publish gate outside the model's prose. **Which MCP publishing approach fits a Git-native blog?** A server that writes repository files and runs deterministic validation is the clearest fit for a Git-native blog. ## Sources - [Tools - Model Context Protocol](https://modelcontextprotocol.io/docs/concepts/tools) — Model Context Protocol - [Model Context Protocol Specification](https://modelcontextprotocol.io/specification/2025-06-18) — Model Context Protocol - [MCP Servers](https://github.com/modelcontextprotocol/servers) — Model Context Protocol - [Ghost Admin API](https://docs.ghost.org/admin-api/) — Ghost - [WordPress REST API Handbook](https://developer.wordpress.org/rest-api/) — WordPress.org - [Strand CMS](https://github.com/BowTiedSwan/strand) — BowTiedSwan --- # AI Editor Agent Profile: One Agent Owns a Publication *By The Strand CMS Team · 2026-08-10 · 9 min read* Canonical: https://www.strandcms.com/blog/hermes-editor-profile > **Summary:** An AI editor agent needs more than a model: it needs a stable profile, editorial policy, source discipline, scheduled phases, and independent publication gates. An **AI editor agent** becomes useful when its ownership is concrete. “Own the blog” is vague. “Research the scheduled batch, write only in this repository, cite every factual claim, run the gates, and stop on failure” is an operating boundary. Hermes Agent's documentation is the authoritative source for Hermes configuration, profiles, skills, tools, and scheduled operation. This article uses that concept as an editorial pattern, not as a claim that configuration alone makes autonomous publication safe. It belongs to the [best MCP server publishing](/blog/best-mcp-for-publishing) pillar and links to [agent publishing safety](/blog/agent-publishing-safety). Product context: [Introducing Strand](/blog/introducing-strand). ## AI editor agent: what the profile owns A durable profile should answer seven questions: - Which repository and directories are in scope? - Which editorial policy applies? - Which skills or procedures are loaded? - Which tools can read, write, validate, or publish? - Which branch and commit conventions apply? - When does each scheduled phase run? - What failures require a hard stop? The profile is not a second CMS. It is the operational memory around the CMS. ## Separate phases, even when one agent runs them A publication job can have research, drafting, QA, and publish phases. Keeping those names separate makes the run easier to resume and audit. The research phase produces a source note. Drafting turns that evidence into articles. QA checks the complete cluster. Publish is the only phase allowed to move the validated batch to production. The agent may perform all four in one day, but the boundaries should remain visible in the logs and in Git history. ## Give the agent a stop condition A profile without stop conditions rewards completion. Add explicit blockers: - working tree is dirty before the run - source URL cannot be retrieved - a claim is not supported - a title or frontmatter field fails schema validation - an internal link is missing - the cluster is incomplete - the branch is not a fast-forward of current main - the publication command fails Stopping is a successful behavior when the alternative is a fabricated or partial release. ## Skills are procedures, not vibes A skill should describe a repeatable method: how to extract claims, what sources to prefer, how to format frontmatter, which commands to run, and what evidence is required before publication. It should be short enough to apply and specific enough to prevent the same mistake twice. The profile can load skills for editorial policy, schema, fact checking, humanization, and SEO/GEO. Those skills should complement one another rather than become a ceremonial checklist. The deterministic commands remain the final test. ## The profile needs least privilege The AI editor agent should not have every tool simply because every tool exists. Read-only research tools, repository write access, validation commands, and production publication credentials can be separate capabilities. If the job cannot perform a destructive action, do not grant it one. MCP helps here by making the available functions explicit, but explicit availability is not enough. Credentials and repository policy still matter. ## A profile review checklist Before scheduling an autonomous editor, verify: - the project path is absolute and correct - the source-of-truth files are checked in - the author and schema exist - the target branch is named - publication mode is known - source and correction policy is written down - the agent can see validation failures - the job reports what changed - the repository can recover from a bad commit ## Ownership is not authority over truth An agent may own the mechanics of a publication without owning reality. It can organize sources, draft prose, run checks, and prepare a release. It cannot promote an unverified number into fact because the sentence sounds plausible. That distinction is especially important for competitive content. A fair editor describes what a competitor does well, cites the documentation, and states where the comparison is uncertain. ## The daily operating loop A scheduled editor can use a simple loop. At the start, inspect the branch and working tree. Next, resolve the date and batch spec from the checked-in calendar. During research, write the source note. During drafting, create the pillar and supports. During QA, run the exact commands and fix failures. During publish, re-check the branch and release only the complete batch. The loop should be observable at each boundary. A later phase should be able to tell whether the earlier phase completed, rather than inferring completion from a half-written file. If the research note is missing, drafting should stop. If a support is missing, publishing should stop. ## Editorial memory versus personal memory A publication profile needs stable project facts: the repository path, schema contract, content calendar, money pages, and source policy. It should not quietly accumulate temporary task progress or unverified beliefs. Temporary state belongs in the run output and Git history; reusable procedures belong in skills or checked-in documentation. That separation keeps the agent from treating yesterday's assumption as today's truth. It also makes the profile easier to audit when the site, package version, or publication policy changes. ## Compare an editor profile with a generic coding agent A generic coding agent can edit a content file. An editor profile adds the why and the stop conditions: which claims need sources, which competitors must be treated fairly, which frontmatter fields are required, and which command is the publication gate. Claude Code's documentation is a useful contrast because it describes a general agentic coding workflow. The profile is the layer that narrows that general capability to one publication's operating rules. The repository then provides the durable source and deterministic implementation. ## Recovering from a bad run When a phase fails, leave the failure legible. Keep the branch, record the command and error, and fix the cause rather than deleting evidence. If a draft is wrong, amend it through a normal commit. If the branch is stale, update it safely. If a source disappears, replace the claim or remove it. A profile that knows how to stop but not how to recover is only half designed. Recovery is what lets the next scheduled run resume without guessing what happened. ## A profile should define the reader Ownership is also editorial focus. This publication writes for technical founders, programmatic-content operators, and developers evaluating publishing infrastructure. That audience needs concrete workflows, current documentation, and fair comparisons. It does not need generic claims that every CMS is “revolutionizing content.” A profile can encode this audience and the desired voice, then leave room for the article brief to define the specific intent. Stable voice is useful; a rigid template that makes every article sound identical is not. ## Scheduling is not a reason to skip gates A cron wake can start a phase, but the clock should not override repository state. If a job wakes late, it still needs to resolve the intended batch and run the same checks. If the branch is dirty or the main branch advanced, the safe choice is to stop or update safely, not to publish stale work. The schedule provides repeatability. It does not provide evidence. Evidence comes from the source note, file contents, command output, and branch history. ## Measure operations, not imagined autonomy Useful metrics include time from research to validated draft, number of failed checks caught before publication, source replacement rate, correction rate, and review time. Avoid claiming that an agent is autonomous because it completed a run. Completion says nothing about accuracy or usefulness. The profile should make those operational facts easy to report. That turns the editor into a system the team can improve rather than a persona everyone is asked to trust. ## A profile should define the handoff Scheduled phases are easier to operate when each one leaves a clear handoff. Research hands over a dated source note. Drafting hands over files with complete frontmatter. QA hands over command output and a clean link graph. Publish hands over the target branch and a release result. The handoff does not need to be elaborate. It needs to be explicit enough that a later wake can distinguish “not started,” “in progress,” “blocked,” and “complete.” That prevents a missed schedule from turning into an accidental partial release. ## Profile drift is a real maintenance problem The publication changes. New content types appear, URLs move, packages are upgraded, and source policies become stricter. Review the profile when those contracts change. Remove obsolete examples and update references rather than letting a stale instruction compete with the checked-in implementation. A stable profile is not a frozen profile. It is a profile whose changes are deliberate, reviewable, and tied to a real project contract. ## Keep the profile small enough to understand An editor profile can become a second operating system if every exception is added permanently. Keep the core policy short, link to the checked-in source-of-truth files, and retire procedures that no longer apply. A smaller profile is easier to reload, review, and correct. ## A profile should document corrections When a published claim changes, the profile should point the editor to the correction path: update the source note, revise the sentence, keep the updated date accurate, and preserve the reason for the change. This is especially important for pricing, product features, and protocol documentation that can move over time. Corrections should be treated as normal editorial work, not as evidence that automation must be abandoned. The useful question is whether the system makes the correction fast, visible, and reviewable. A profile that preserves those habits is more valuable than one that merely claims autonomy. That is the standard worth carrying into the next scheduled phase. The profile should make that standard visible and easy to apply. This keeps the process practical. ## Why this matters An editor profile can become a second operating system if every exception is added permanently. Keep the core policy short, link to the checked-in source-of-truth files, and retire procedures that no longer apply. A smaller profile is easier to reload, review, and correct. ## A profile should document corrections When a published claim changes, the profile should point the editor to the correction path: update the source note, revise the sentence, keep the updated date accurate, and preserve the reason for the change. This is especially important for pricing, product features, and protocol documentation that can move over time. Corrections should be treated as normal editorial work, not as evidence that automation must be abandoned. The useful question is whether the system makes the correction fast, visible, and reviewable. A profile that preserves those habits is more valuable than one that merely claims autonomy. That is the standard worth carrying into the next scheduled phase. The profile should make that standard visible and easy to apply. This keeps the process practical. ## Why this matters A stable editor profile turns an irregular prompt into a repeatable publication operation. The gain is not magical autonomy. The gain is fewer hidden decisions: the agent knows its scope, the repository shows its work, and the release command can refuse an incomplete or failing batch. ## FAQ **What is an AI editor agent?** It is an agent configured to perform defined editorial work, such as research, drafting, QA, and release, within a bounded publication workflow. **What belongs in an editor profile?** The profile should define the project scope, skills, source policy, tool permissions, schedule, branch rules, and the conditions that block publication. **Can an editor profile replace editorial standards?** No. It can operationalize standards, but sources, claims, fairness, validation, and correction rules still need to be explicit. ## Sources - [Hermes Agent documentation](https://hermes-agent.nousresearch.com/docs) — Nous Research - [Claude Code overview](https://docs.anthropic.com/en/docs/claude-code/overview) — Anthropic - [Building effective agents](https://www.anthropic.com/research/building-effective-agents) — Anthropic - [About rulesets](https://docs.github.com/en/repositories/configuring-branches-and-merges-in-your-repository/managing-rulesets/about-rulesets) — GitHub - [Strand CMS](https://github.com/BowTiedSwan/strand) — BowTiedSwan --- # Claude Code Blog Publishing in 2026: Full Pipeline *By The Strand CMS Team · 2026-08-10 · 9 min read* Canonical: https://www.strandcms.com/blog/publish-blog-claude-code > **Summary:** A Claude Code blog workflow should separate research, drafting, validation, review, and publication. The model can accelerate the work without owning the final gate. A **Claude Code blog** workflow works best when Claude Code is the operator for a repository, not the unreviewed owner of production. Anthropic documents Claude Code as an agentic coding tool that works in a codebase, and its MCP documentation explains how it can connect to external tools. Those capabilities are useful only when the publication system supplies the boundaries. This guide belongs to the [best MCP server publishing](/blog/best-mcp-for-publishing) pillar. It also connects to [agent publishing safety](/blog/agent-publishing-safety) and [the anatomy of a publishing MCP](/blog/anatomy-publishing-mcp). If you want the product example, read [Introducing Strand](/blog/introducing-strand). ## Claude Code blog pipeline Use six stages: 1. Define the article and its sources. 2. Create a branch from the current main branch. 3. Draft the MDX and frontmatter. 4. Run deterministic validation and batch QA. 5. Review the diff and link graph. 6. Publish only after the complete batch passes. The important design choice is that the model can propose changes at every stage, while the final state is checked by scripts and repository policy. ## 1. Start with a source brief Write down the target keyword, search intent, audience, article type, internal links, and source URLs before asking Claude Code to draft. For factual topics, record what each source supports. A source list is not proof that every sentence is correct; it is a map for checking claims. Prefer primary documentation for product capabilities. If a page says an API exists, cite the vendor's API docs. If it says a protocol defines tools, cite the protocol documentation. If you cannot verify a claim, cut it or label it as analysis. ## 2. Give Claude Code a narrow repository task Claude Code is more reliable when the task names the files it may touch and the checks it must run. A useful brief says: - work only on `content/batch-3` - create the pillar first - use the supplied date for every article - begin every body with `` - keep sources visible in frontmatter and the body - run `npm run validate` - run `npm run content:qa -- --batch=3` - stop if a gate fails This is not bureaucracy. It reduces the number of hidden assumptions in the prompt and makes the resulting diff easier to review. ## 3. Draft the pillar before the supports The pillar defines the comparison frame, vocabulary, and internal-link structure. Write it first, then link every support back to it. The pillar should link every support so the cluster is navigable from both directions. For this cluster, the pillar is [Best MCP Server Publishing Tools in 2026](/blog/best-mcp-for-publishing). Its supports cover Claude Code, quality guardrails, safety, agent ownership, and the tool lifecycle. That is a useful cluster because it moves from choosing a path to operating one. ## 4. Keep frontmatter machine-checkable A blog file needs more than prose. Set the title, slug, description, dates, author, status, content type, primary keyword, keywords, summary, FAQ, and sources. Keep the slug identical to the filename. Keep the description within the schema's length limits. The `status: published` field can be present before the batch publish phase in this workflow, but it is not a substitute for the publish gate. It expresses the intended final state; the batch command is what performs the direct publication. ## 5. Use MCP for bounded operations The MCP tools documentation describes discoverable tools with structured input schemas. That maps cleanly to editorial operations: ```text create_draft(slug, frontmatter, body) validate_post(slug) content_qa(batch) publish_batch(batch) ``` The names are less important than the boundaries. A draft operation should not also push to production. A validation result should include actionable errors. A publish operation should verify the complete batch instead of trusting the model's claim that all files exist. ## 6. Run the checks outside the model Run the schema validator and batch QA in the shell. Read the output. Fix the actual error. Repeat until both commands pass. Do not ask Claude Code to “assume validation passed,” and do not replace a failed check with a prose explanation. Useful checks include: - required frontmatter and valid author - title, slug, and description constraints - `` as the first body component - keyword placement without stuffing - word-band compliance - FAQ and source counts - a link to the pillar and a money page - the pillar linking every support ## 7. Review the diff like an editor Read the rendered shape, not just the beginning of the file. Look for claims that outrun the citations, vague competitor descriptions, repeated paragraphs, and links that point to the wrong slug. Confirm every comparison names at least one genuine competitor strength. For an AI content agent, the review should also ask whether the article sounds like a real editor wrote it. Remove empty transitions, generic “in today's landscape” openings, and claims that are impressive but uncheckable. ## 8. Publish the complete cluster A safe batch does not publish one article because that one article happens to pass. Confirm all planned files exist, all intended statuses are correct, validation passes, and batch QA passes. Then use the repository's batch publication command. Branch protection and rulesets can add another independent control. GitHub documents rulesets as a way to define repository rules and enforcement. Use them to require checks where your deployment model supports that arrangement. ## Troubleshooting common failures **The keyword check fails.** Put the exact primary keyword in the title, first section, an H2, and the description, then read the copy to make sure it still sounds natural. **The link graph fails.** Add the pillar link to each support. Add every support slug to the pillar. Include a money-page link such as `/` or the Strand introduction. **A source does not support the sentence.** Rewrite the sentence to match what the source actually says, attribute it, or remove it. Never widen a source's claim to save a draft. **The publish command rejects the batch.** Treat that as a blocker. Inspect status, branch ancestry, missing files, and gate output. Do not force-push. ## A concrete repository layout A small repository can make the pipeline obvious: ```text content/posts/ # article source content/authors/ # author records automation/research/ # source notes scripts/validate.ts # schema gate scripts/qa-content-batch.mjs # cluster gate ``` The exact layout can differ, but the principle is useful: keep source, research, and checks close enough that a reviewer can navigate from a claim to its evidence and from a post to the command that validates it. ## Research first, prose second Ask Claude Code to collect and summarize source evidence before it writes the article. The research pass should identify which claims are supported, where sources disagree, and which planned claims need to be softened. This prevents the common failure mode in which a polished draft causes the agent to retrofit citations after the fact. For vendor documentation, capture the page title, URL, publisher, and the narrow claim it supports. For technical guidance, retain the relevant section in the research note so a later editor can re-check the interpretation. ## Keep publication permissions separate Claude Code can connect to MCP servers, but access should be deliberate. Use read-only tools for research where possible. Let draft operations write only to the working branch. Make publication require a specific batch and a fresh validation run. If the client cannot express this separation, enforce it in the server or shell command. A model that can edit a file does not automatically need permission to push the branch. Fewer permissions mean fewer ways for a mistaken instruction to become a production incident. ## Review the generated page, not just MDX MDX can look correct while rendering a broken link, an unexpected component, or metadata that is not emitted where the reader and crawler can see it. Run the site's normal build or preview when the change includes a component, table, or unusual Markdown. Read the first screen as a reader: is the answer clear, are sources visible, and is the next action obvious? ## A final preflight checklist Before asking Claude Code to publish, check the repository state and the target date. Confirm the branch is the intended batch branch and that the working tree contains no unrelated edits. Confirm every article in the calendar has a matching file and that the pillar-first link graph is complete. Read the first paragraph of every article. It should answer the query without requiring the reader to know the product. Then scan all claims that mention a vendor, a version, a protocol feature, a price, or a number. Each should have a source or a clear attribution. If a sentence is merely advice, write it as advice rather than dressing it up as a documented fact. Run the validator after every meaningful frontmatter change. Run batch QA after all links are present. If a check fails, keep the failure visible while fixing it. Do not edit the check or remove a required field just to make the output green. ## A note on generated content Claude Code can produce a useful first draft quickly, but the final article should not read like a transcript of the agent. Replace generic transitions with direct claims. Keep examples concrete. Remove duplicated conclusions. Preserve caveats where the evidence is limited. The goal is not to hide that software helped; the goal is to publish something a reader can verify and use. ## A small worked sequence A practical run starts with a source note containing the target docs and the claims they support. Claude Code creates the pillar, then creates the support articles with the same date and author. The validator catches frontmatter mistakes. Batch QA catches a missing support link. The editor reads the corrected diff, and only then does the publication command run. Each step leaves a state that the next step can inspect. The same sequence works for a single post when a complete cluster is not needed. The principle stays the same: the agent writes a proposed state, the checks inspect the actual state, and publication is an explicit side effect rather than an implied ending. That separation gives the editor a clear place to stop. It also makes a failed run easier to resume because the last durable state is visible. The workflow stays understandable when the tool list grows. That is the real benefit of a pipeline. It keeps the tool list understandable when the publication grows and gives the editor a stable place to inspect the result. It is a small operational advantage, but a durable one. ## Why this matters Claude Code can make repository-based publishing faster, but speed is not the editorial standard. The standard is a complete, sourced, readable cluster whose final state can be checked without trusting the model's self-report. The standard is a complete, sourced, readable cluster whose final state can be checked without trusting the model's self-report. The practical formula is simple: let the agent handle mechanical work, let deterministic checks reject malformed work, and let a human-readable diff preserve accountability. ## FAQ **Can Claude Code publish a blog post?** It can work with repository files and connected MCP tools, but a safe workflow keeps publication behind validation, review, and explicit permission. **What should Claude Code do before publishing?** It should verify sources, write the required frontmatter, run schema validation and batch QA, inspect the diff, and confirm the complete cluster is present. **Is Claude Code a CMS?** No. Claude Code is an agentic coding tool; the CMS, repository, validation scripts, and deployment process still provide the publication system. ## Sources - [Claude Code overview](https://docs.anthropic.com/en/docs/claude-code/overview) — Anthropic - [Claude Code MCP](https://docs.anthropic.com/en/docs/claude-code/mcp) — Anthropic - [Tools - Model Context Protocol](https://modelcontextprotocol.io/docs/concepts/tools) — Model Context Protocol - [About rulesets](https://docs.github.com/en/repositories/configuring-branches-and-merges-in-your-repository/managing-rulesets/about-rulesets) — GitHub - [Strand CMS](https://github.com/BowTiedSwan/strand) — BowTiedSwan --- # llms.txt Adoption in 2026: A 100-Site Snapshot *By The Strand CMS Team · 2026-08-10 · 9 min read* Canonical: https://www.strandcms.com/blog/llms-txt-adoption-data > **Summary:** This llms.txt adoption snapshot checks 100 named domains on August 10, 2026. It measures endpoint responses for the sample, not web-wide adoption, crawler use, rankings, or citations. This llms.txt adoption report is a narrow endpoint-status snapshot, not a census of the web. The phrase llms.txt adoption here means endpoint responses in this named sample, not confirmed use by an AI engine. On August 10, 2026, we requested `https:///llms.txt` for a fixed list of 100 named domains. The result tells us which endpoints returned `200` in that run. It does not tell us whether an AI engine consumed a file, whether a site owner endorses the proposal, or whether the file changes rankings or citations. The [complete llms.txt guide](/blog/llms-txt-complete-guide) provides the technical context. ## llms.txt adoption: method first The sample was selected before collection across AI, search, developer tooling, CMS, infrastructure, SaaS, and documentation/media categories. It is a convenience sample designed to make the raw checks reproducible, not a statistically representative frame of all websites. For each domain we sent: ```text GET https:///llms.txt User-Agent: StrandResearch/1.0 Timeout: 12 seconds Follow redirects: yes ``` We classified a final `200` as “found for this test.” We recorded `403`, `404`, `429`, and transport errors separately. A `200` means the endpoint responded successfully. It does not establish that the response is a valid proposal file, that a crawler can access every linked page, or that a model uses it. ## The 100-domain raw record The complete record below is the dataset used in this article. Keeping the raw statuses visible matters more than presenting one attractive percentage. | # | Domain | Status | | ---: | --- | --- | | 1 | openai.com | 403 | | 2 | anthropic.com | 404 | | 3 | perplexity.ai | 403 | | 4 | google.com | 404 | | 5 | developers.google.com | 404 | | 6 | cloud.google.com | 404 | | 7 | microsoft.com | 404 | | 8 | bing.com | 404 | | 9 | github.com | 200 | | 10 | gitlab.com | 403 | | 11 | bitbucket.org | 404 | | 12 | stackoverflow.com | 404 | | 13 | npmjs.com | 403 | | 14 | pypi.org | 404 | | 15 | docker.com | 404 | | 16 | kubernetes.io | 404 | | 17 | cncf.io | 404 | | 18 | linuxfoundation.org | 404 | | 19 | mozilla.org | 404 | | 20 | developer.mozilla.org | 404 | | 21 | web.dev | 404 | | 22 | w3.org | 404 | | 23 | ietf.org | 404 | | 24 | rfc-editor.org | 404 | | 25 | wordpress.org | 200 | | 26 | ghost.org | 404 | | 27 | sanity.io | 200 | | 28 | contentful.com | 429 | | 29 | strapi.io | 403 | | 30 | payloadcms.com | 200 | | 31 | hashnode.com | 200 | | 32 | beehiiv.com | 403 | | 33 | substack.com | 404 | | 34 | medium.com | 403 | | 35 | dev.to | 404 | | 36 | vercel.com | 200 | | 37 | netlify.com | 200 | | 38 | astro.build | 404 | | 39 | nextjs.org | 200 | | 40 | react.dev | 200 | | 41 | vuejs.org | 200 | | 42 | svelte.dev | 200 | | 43 | angular.dev | 200 | | 44 | deno.com | 200 | | 45 | bun.sh | 200 | | 46 | nodejs.org | 200 | | 47 | python.org | 404 | | 48 | rust-lang.org | 404 | | 49 | go.dev | 404 | | 50 | php.net | 404 | | 51 | ruby-lang.org | 200 | | 52 | rails.org | 000/ERR | | 53 | django-project.com | 000/ERR | | 54 | fastapi.tiangolo.com | 404 | | 55 | huggingface.co | 404 | | 56 | arxiv.org | 404 | | 57 | kaggle.com | 200 | | 58 | databricks.com | 200 | | 59 | snowflake.com | 404 | | 60 | tableau.com | 404 | | 61 | segment.com | 200 | | 62 | mixpanel.com | 404 | | 63 | amplitude.com | 200 | | 64 | stripe.com | 200 | | 65 | shopify.com | 200 | | 66 | slack.com | 200 | | 67 | notion.so | 200 | | 68 | linear.app | 200 | | 69 | figma.com | 404 | | 70 | atlassian.com | 200 | | 71 | trello.com | 200 | | 72 | asana.com | 200 | | 73 | monday.com | 200 | | 74 | clickup.com | 200 | | 75 | zapier.com | 200 | | 76 | hubspot.com | 200 | | 77 | salesforce.com | 200 | | 78 | intercom.com | 200 | | 79 | drift.com | 404 | | 80 | mailchimp.com | 200 | | 81 | ahrefs.com | 404 | | 82 | semrush.com | 200 | | 83 | moz.com | 404 | | 84 | searchenginejournal.com | 200 | | 85 | searchengineland.com | 200 | | 86 | search.google.com | 404 | | 87 | blog.google | 404 | | 88 | openai.github.io | 404 | | 89 | docs.github.com | 200 | | 90 | docs.docker.com | 200 | | 91 | docs.python.org | 404 | | 92 | docs.npmjs.com | 404 | | 93 | docs.astro.build | 404 | | 94 | docs.ghost.org | 200 | | 95 | docs.sanity.io | 200 | | 96 | docs.strapi.io | 200 | | 97 | docs.payloadcms.com | 000/ERR | | 98 | docs.contentful.com | 429 | | 99 | readthedocs.org | 404 | | 100 | git-scm.com | 404 | ## What the snapshot can say The record shows a mixture of successful responses, missing responses, access denials, rate limits, and transport failures. That is the useful finding: endpoint status is not a single binary measure of adoption. If we calculate a found rate, it must use the declared rule—final `200` only—and remain explicitly limited to this named sample and collection date. The raw table also prevents a misleading conclusion from a single count. A `403` can mean the server refused this request; a `429` can mean rate limiting; `000/ERR` means the run did not obtain a normal HTTP response. None should be silently converted into “the site does not have llms.txt.” ## What the snapshot cannot say It cannot establish web-wide llms.txt adoption. It cannot establish that a `200` body follows the proposal or that a model fetches it. It cannot show whether a site appears in an answer, whether a citation is accurate, or whether the file affects Google Search. OpenAI's crawler documentation and Perplexity's crawler documentation describe vendor behavior, but endpoint presence is not evidence of consumption. It also cannot justify ranking the sampled companies by AI readiness. The domains were selected for coverage, not randomly sampled, and several belong to the same organizations or documentation ecosystems. ## How to reproduce the check Save the domain list, run one request per domain with the declared user agent and timeout, record the final status and redirect chain, then publish the raw record with the collection date. Do not change the sample midway. If you re-run it, label the result as a new snapshot rather than silently editing history. A responsible follow-up would validate whether successful bodies are parseable Markdown and whether their links are public and useful. That is a different measurement from endpoint status and should be reported separately. ## Why this matters Data about emerging web conventions is easy to inflate. A 100-site check becomes useful when the sample, request, classification, raw record, and limitations are visible. It becomes misleading when one endpoint response is presented as proof of crawler behavior or search performance. For implementation, start with the [llms.txt guide](/blog/llms-txt-complete-guide), then compare [llms.txt, llms-full.txt, and Markdown twins](/blog/llms-txt-vs-llms-full). For a publishing system that generates machine-readable artifacts from Git-backed content, see [Introducing Strand CMS](/blog/introducing-strand). ## Reading a snapshot without overfitting A sample like this is most useful as a baseline for future work. If the same 100 domains are checked again, the comparison can show how endpoint statuses changed under the same method. That still will not become a representative adoption estimate, but it can reveal whether the convention is becoming more common in this selected group. The next measurement should add body validation: check whether a successful response is Markdown, whether it has a useful title and description, and whether its links resolve. That would answer a different question from “did the endpoint return 200?” Keeping the questions separate is how a small research note avoids turning into AI slop. The raw record also makes corrections possible. If a domain was mistyped, a redirect was mishandled, or the collection date was wrong, readers can identify the affected row. Future runs should preserve the original record and publish a new dated result rather than rewriting history. ## What a follow-up study should add A second run should preserve the same 100 domains and method before expanding the sample. It could add a body-shape check: content type, Markdown headings, link count, and whether the body identifies a canonical site. A third layer could record redirects and response dates. Each layer should have its own pass or fail definition. It would also be useful to publish the collection script and a machine-readable CSV. Reproducibility is stronger when another researcher can run the same request with the same timeout and compare the raw output. The result will still be a snapshot, but it will be a more inspectable one. The study should not silently mix endpoint presence with AI visibility. To investigate citations, you would need a separate query set, a dated retrieval protocol, and a way to distinguish a citation from a model's unsupported paraphrase. That is a different research project with different sources and confounders. ## Editorial interpretation The responsible conclusion is modest: some selected domains returned a successful response, many did not, and several could not be classified as absent because access or transport intervened. The result is enough to justify better measurement; it is not enough to declare a standard adopted or rejected. That modesty is useful for operators. If you publish the endpoint, describe it accurately, link canonical pages, and watch for stale content. If you do not publish it, keep your HTML and source links strong. The snapshot does not turn either choice into a guaranteed search outcome. ## FAQ ### Why report the raw table? The raw table lets readers audit the collection and see that 403, 429, 404, and transport errors were not treated as the same event. ### Is a successful endpoint enough for AI visibility? No. It is only one observable property. Crawl access, page quality, relevance, and a model's retrieval and citation choices remain separate. ### When should this dataset be updated? Re-run it when you want a new dated snapshot, preserve the old method, and record any sample or request changes before collection. ## Sources - [The /llms.txt file](https://llmstxt.org/) - [Answer.AI llms-txt repository](https://github.com/AnswerDotAI/llms-txt) - [OpenAI crawler documentation](https://developers.openai.com/api/docs/bots) - [Perplexity crawler documentation](https://docs.perplexity.ai/docs/resources/perplexity-crawlers) - [Google robots.txt specification](https://developers.google.com/search/docs/crawling-indexing/robots/robots_txt) - [Google helpful content guidance](https://developers.google.com/search/docs/fundamentals/creating-helpful-content) ## FAQ **What does this llms.txt adoption study measure?** It measures the final HTTP response category from a GET request to /llms.txt on 100 named domains collected on August 10, 2026. **Does a 200 response prove a site uses llms.txt?** It proves that the endpoint returned a successful response during this test. It does not prove that an AI engine reads or uses the file. **Is this a representative estimate of llms.txt adoption?** No. It is a fixed, category-balanced snapshot of 100 selected domains and should not be generalized to all websites. **Why separate 403 and 429 from 404?** 403 and 429 can indicate access restrictions or rate limits, so they are not equivalent to a confirmed missing endpoint. ## Sources - [The /llms.txt file](https://llmstxt.org/) — llms-txt - [llms-txt reference repository](https://github.com/AnswerDotAI/llms-txt) — Answer.AI - [Overview of OpenAI Crawlers](https://developers.openai.com/api/docs/bots) — OpenAI - [Perplexity Crawlers](https://docs.perplexity.ai/docs/resources/perplexity-crawlers) — Perplexity - [robots.txt Specifications](https://developers.google.com/search/docs/crawling-indexing/robots/robots_txt) — Google Search Central - [Creating Helpful, Reliable, People-First Content](https://developers.google.com/search/docs/fundamentals/creating-helpful-content) — Google Search Central --- # The Complete llms.txt Guide for 2026 *By The Strand CMS Team · 2026-08-10 · 9 min read* Canonical: https://www.strandcms.com/blog/llms-txt-complete-guide > **Summary:** This llms.txt guide explains the proposed Markdown index for important site resources. It is a convention for machine-readable orientation, not a proven ranking signal or citation guarantee. An llms.txt file is a proposed Markdown index for a website. This llms.txt guide explains how it gives a language-model tool a short explanation of the site and links to the pages that matter most. This [llms.txt guide](/blog/llms-txt-complete-guide) is the pillar for our technical authority cluster; it explains the convention without pretending that presence proves usage, ranking impact, or AI citations. The practical answer is simple: publish a small, accurate file at `/llms.txt`, keep it current, and treat it as an additional representation of your information architecture. Do not replace normal crawlable pages, `robots.txt`, internal links, or editorial quality with it. ## llms.txt guide: what the proposal is The [llms-txt proposal](https://llmstxt.org/) describes a root-level Markdown file with a short overview and grouped links. The suggested shape is intentionally readable by humans as well as machines. A useful file tells a visitor what the site is, then points to canonical resources such as documentation, product pages, policies, or reference material. That makes `/llms.txt` an orientation layer. It is closer to a carefully edited index than to an access-control file or a new search metadata field. The proposal is not the same thing as a vendor contract. A site can publish the file and an engine can still ignore it. ## What to put in the file A small site can start with this shape: ```md # Example Docs > Canonical documentation for Example, a hosted API. ## Start here - [Overview](https://example.com/docs/) - [Quickstart](https://example.com/docs/quickstart) ## Reference - [API reference](https://example.com/docs/api) - [Authentication](https://example.com/docs/auth) ## Optional - [Changelog](https://example.com/changelog) ``` Use absolute URLs, descriptive link labels, and sections that reflect how a reader would navigate the site. Keep the summary factual. If a page is temporary, private, or not intended for discovery, do not add it merely to make the file longer. The file should point to canonical pages, not duplicate every paragraph. A curated list is easier to review and less likely to drift than a generated dump containing navigation, tracking URLs, or stale routes. ## What llms.txt is not `llms.txt` does not replace `robots.txt`. Google documents `robots.txt` as a crawler-access mechanism with matching rules and limits. An informational Markdown file does not grant access to a page, revoke access, or create a security boundary. Never put secrets in either file; access control belongs in authentication and server policy. It is also not proof that an AI engine reads your site. OpenAI documents several crawler identities and purposes in its [crawler guidance](https://developers.openai.com/api/docs/bots), while Perplexity documents its own crawlers and user-fetch behavior. Neither document says that every site must provide `/llms.txt` for inclusion. Finally, it is not a substitute for a useful page. A model still needs an accessible, understandable destination. The linked page needs a clear answer, stable URL, and evidence appropriate to the claim. ## llms.txt versus llms-full.txt The concise file is useful as a map. Some implementations also expose an expanded `llms-full.txt` containing more page content. That can reduce the number of requests for a consumer that chooses to fetch it, but it creates maintenance and size costs. It also raises the risk of including stale, duplicated, or private material. There is no universal requirement to publish the expanded variant. Choose it when you can generate it from canonical content, exclude drafts, and test that it remains within your operational limits. Otherwise, a precise index plus good pages is the safer starting point. See [llms.txt vs llms-full.txt vs .md pages](/blog/llms-txt-vs-llms-full) for the representation tradeoffs. ## How to publish it safely 1. Decide which public pages are genuinely canonical. 2. Write a one-paragraph site description that does not oversell the product. 3. Group links by reader task, not by internal database type. 4. Serve the file at the root over HTTPS with a successful response. 5. Check links, redirects, and accidental private URLs in CI. 6. Rebuild it whenever your information architecture changes. A generated file is usually safer than hand-editing once the list becomes large, but generation must have an allowlist. “Everything in the database” is not a content strategy. It can expose drafts, account pages, duplicate query URLs, and obsolete documentation. ## How to test it Start with a direct request: ```bash curl -iL https://example.com/llms.txt ``` Confirm the final response is successful, the body is readable Markdown, and the links resolve to the intended canonical pages. Then test the destination pages without relying on client-side clicks. If the answer is only available after a browser executes an application shell, the index has not solved the rendering problem. A directly accessible Markdown representation can complement HTML; it cannot excuse broken HTML. For a page-level implementation pattern, read [Markdown Twins](/blog/markdown-twins). For crawl policy, read [AI crawler robots.txt](/blog/robots-txt-ai-crawlers). For the dated endpoint snapshot, read [llms.txt Adoption Data](/blog/llms-txt-adoption-data). ## Adoption data needs careful language A status check across a named sample can tell you how many endpoints responded during that run. It cannot tell you how many sites use the proposal across the web, whether an engine consumed the file, or whether the file changed citations. Report the date, domains, request method, response categories, and limitations. Separate `404` from `403`, `429`, and transport failures. That distinction matters because access controls and rate limits are not the same as confirmed absence. A responsible report says “not found for this test,” not “the owner does not have the file.” ## Strand's implementation choice Strand keeps content as MDX in Git and publishes machine-readable artifacts alongside the canonical site. Its [repository](https://github.com/BowTiedSwan/strand) is the source for Strand-specific implementation claims. The useful lesson is not that one CMS guarantees AI visibility; it is that a reproducible source and deterministic validation make machine-readable output easier to maintain. See [Introducing Strand CMS](/blog/introducing-strand) for the product overview and the [Markdown Twins](/blog/markdown-twins) implementation guide. ## Why this matters The value of llms.txt is operational clarity, not a magical signal. A concise index gives people and tools a stable starting point, exposes the pages you consider important, and forces an honest conversation about your information architecture. It is cheap to test and easy to remove if it does not fit your stack. The risk is overclaiming. Publishing a file does not guarantee crawling, indexing, ranking, or citation. Keep the file accurate, keep the destination pages useful, and measure only what your own collection or logs can actually support. ## Maintenance checklist for a living file The first publish is the easy part. The file becomes useful only when it remains aligned with the site. Put it next to the content build or route configuration so a change to a canonical page can trigger a review. A weekly or release-based check is enough for many small sites; a frequently changing documentation site may need generation on every deploy. Review the overview for accuracy, then check each URL. Remove redirects that point to retired content, merge duplicate resources, and keep the headings understandable to a person who has never seen the site. Do not add a link merely because it exists in a database. The file is a map, so every link should earn its place. Keep a record of exclusions. If an account area, internal handbook, or draft collection is deliberately absent, that is a content boundary worth preserving. Test the generated artifact with a fixture containing unpublished content. The test should prove that the fixture cannot leak into the public response. A good review also checks the claims in the description. If the product, API, or policy changed, update the summary instead of leaving a stale sentence that a downstream consumer may repeat. A short file is easier to audit, which is one reason not to turn the root endpoint into a complete archive. ## Common implementation mistakes The most common mistake is returning a branded HTML error page with a successful status. A consumer that requests Markdown should get Markdown, not an application shell. The second is using relative links that resolve differently for different consumers. Absolute canonical URLs are less ambiguous in a root-level index. Another mistake is treating a `200` check as proof of a valid file. Validate the content type, body shape, link targets, and public status of each destination. A file can respond successfully while pointing to pages that no longer exist. Finally, do not use the file to hide a weak information architecture. If the important answer cannot be expressed in a clear page title, summary, and heading structure, the index is only documenting the problem. Fix the page first, then expose it. ## Editorial review before release Have one person read the file as a map and another read every destination page. The first reviewer checks whether the grouping reflects the site's public priorities. The second checks whether each link still supports the description and whether the destination is genuinely canonical. This catches a failure that syntax checks cannot: a perfectly valid file that directs readers to the wrong explanation. Check the format with a plain text editor, not only a rendered browser preview. Confirm headings are recognizable, links have useful labels, and no template tokens remain. If the file is generated, inspect the build diff. A change that suddenly adds hundreds of URLs should require an explicit reason. Use a content inventory to decide what belongs in the optional section. Optional does not mean “everything else.” It means material that may help a consumer but is not part of the primary path. If the section grows faster than the main sections, the file is becoming an archive and should be split into better-curated resources. ## A compact operating policy The team can write down four rules: only public canonical pages may appear; every link must resolve; every description must be supportable; and the generated file must be checked before deploy. These rules are simple enough for a solo operator and strong enough to prevent the most damaging leaks. When a page is unpublished, remove it from the index in the same change that removes it from the public build. When a URL moves, update both the page and the file. When a claim becomes time-sensitive, add a date or soften the language. The file should never become an unowned copy of product marketing. ## FAQ ### Can llms.txt control AI crawlers? No. It is an informational convention. Use `robots.txt`, authentication, and server controls for access policy. ### Should every page appear in llms.txt? No. Curate the pages that explain the site and serve the reader's main tasks. A shorter, accurate index is easier to trust than a full URL dump. ### Does the file replace a sitemap? No. A sitemap and llms.txt have different purposes. Keep normal discovery and canonical-link systems working independently. ## Sources - [The /llms.txt file](https://llmstxt.org/) - [Answer.AI llms-txt repository](https://github.com/AnswerDotAI/llms-txt) - [Google robots.txt specification](https://developers.google.com/search/docs/crawling-indexing/robots/robots_txt) - [OpenAI crawler documentation](https://developers.openai.com/api/docs/bots) - [Perplexity crawler documentation](https://docs.perplexity.ai/docs/resources/perplexity-crawlers) - [Strand CMS repository](https://github.com/BowTiedSwan/strand) ## FAQ **What is llms.txt?** llms.txt is a proposed Markdown file, usually placed at a site's root, that summarizes important resources and links to them for language-model tools. **Does llms.txt improve Google rankings?** There is no established evidence that publishing llms.txt improves Google rankings. Treat it as a machine-readable orientation layer, not a ranking shortcut. **Where should llms.txt live?** The convention places it at the root of the site, at https://example.com/llms.txt, so the location is stable and easy to test. **Is llms-full.txt required?** No. An expanded file can be useful, but the proposal does not make it a universal requirement and engines may ignore either file. ## Sources - [The /llms.txt file](https://llmstxt.org/) — llms-txt - [llms-txt reference repository](https://github.com/AnswerDotAI/llms-txt) — Answer.AI - [robots.txt Specifications](https://developers.google.com/search/docs/crawling-indexing/robots/robots_txt) — Google Search Central - [Overview of OpenAI Crawlers](https://developers.openai.com/api/docs/bots) — OpenAI - [Perplexity Crawlers](https://docs.perplexity.ai/docs/resources/perplexity-crawlers) — Perplexity - [Strand CMS on GitHub](https://github.com/BowTiedSwan/strand) — BowTiedSwan --- # llms.txt vs llms-full.txt vs .md Pages *By The Strand CMS Team · 2026-08-10 · 7 min read* Canonical: https://www.strandcms.com/blog/llms-txt-vs-llms-full > **Summary:** llms.txt is a concise index, llms-full.txt is an optional expanded document, and .md pages are page-level representations. None is a universal instruction channel or citation guarantee. `llms.txt` is a short index, `llms-full.txt` is an optional expanded representation, and a `.md` page is a page-level Markdown twin. The right choice depends on what you need to expose and maintain. None of these formats is a universal command that every AI engine must read. This article supports the [llms.txt guide](/blog/llms-txt-complete-guide) with a narrower implementation comparison. The safe order is: make the canonical HTML correct, publish a concise index if it helps, then add Markdown representations when you can guarantee parity and privacy. ## llms-full.txt: the expanded option The proposal at [llmstxt.org](https://llmstxt.org/) centers on a concise file that explains a site and links to useful resources. An expanded file can place more of that linked material into one response. That may be convenient for a consumer that chooses to fetch it, but it also creates a second content surface. The main engineering question is not “can we concatenate pages?” It is “can we keep the concatenation accurate?” A full dump can include stale headings, duplicate navigation, tracking links, draft text, or content that was meant to be behind a permission boundary. If you generate it, use the same canonical source as the HTML build and an explicit public-content allowlist. ## `.md` pages: a page-level representation A Markdown twin is a clean text representation of one canonical page. It can be exposed at a predictable route such as `/docs/intro.md`, through content negotiation, or through a separate endpoint. [MDN's content-negotiation guide](https://developer.mozilla.org/en-US/docs/Web/HTTP/Content_negotiation) and [RFC 9110](https://www.rfc-editor.org/rfc/rfc9110) describe the broader HTTP idea: a resource can have multiple representations, and clients and servers need clear semantics for choosing and caching them. A `.md` route is not automatically a protocol. Do not claim that every crawler requests it. The benefit is practical: a machine or developer can retrieve the article without parsing a visual shell, while the canonical HTML remains the user-facing page. ## What AI engines actually read The honest answer is that public vendor documentation describes crawler identities and access behavior, not a universal preference for one filename. An engine may fetch HTML, use a search index, follow links, or retrieve a page at a user's request. Its behavior can change. That is why format claims need boundaries. You can verify that an endpoint exists, returns the expected representation, and contains the same public answer as the canonical page. You cannot infer from a successful `curl` that a model consumed it or cited it. Google's [JavaScript SEO guidance](https://developers.google.com/search/docs/crawling-indexing/javascript/javascript-seo-basics) is useful here: important content should be available in output that can be crawled and rendered. A Markdown twin may reduce parsing work, but it does not remove the need for crawlability, links, and useful content. ## A comparison table | Representation | Primary job | Main strength | Main risk | Use it when | | --- | --- | --- | --- | --- | | `llms.txt` | Curated site index | Small and easy to review | Stale links or overclaiming | You need a machine-readable map | | `llms-full.txt` | Expanded site bundle | Fewer fetches for a willing consumer | Duplication, size, privacy leaks | You can generate and validate it | | `.md` page | One-page text representation | Clear page-level retrieval | Parity drift or false protocol claims | You can keep HTML and Markdown aligned | | HTML | Canonical reader page | Broad compatibility and semantics | Client-only content or noisy templates | Always; this remains the foundation | ## Discovery, access, rendering, citation These are four different problems: 1. **Discovery:** can a consumer find the URL through links, search, or an index? 2. **Access:** does the server allow the request and return the intended status? 3. **Rendering:** is the important content present in a usable representation? 4. **Citation:** does an engine decide that the source is relevant and trustworthy enough to quote? An llms.txt file mostly addresses discovery. A Markdown twin mostly addresses representation. `robots.txt` addresses crawler policy, within its documented limits. None of them guarantees the fourth step. The [AI crawler robots.txt](/blog/robots-txt-ai-crawlers) guide covers access policy separately. Keeping these layers separate prevents a common mistake: treating a file format as proof of search performance. ## How to build a Markdown twin Start with one route and one content type. Render the title, summary, publication metadata, headings, lists, tables, and source links. Strip navigation, comments, account controls, and decorative UI. Preserve the canonical URL prominently so a reader can move between representations. Then test parity: - the same public title appears in both outputs - the answer paragraph exists in both outputs - links resolve to the same destinations - no draft or private field appears in Markdown - dates and updated state match - the HTML page remains canonical - cache headers do not serve stale text after a publish A route that is generated from the same MDX source as HTML has a simpler parity story. Strand's [GitHub repository](https://github.com/BowTiedSwan/strand) is the source for its own MDX and generated-artifact claims; do not generalize that implementation to every CMS. ## When a full dump is a bad idea Do not add `llms-full.txt` if the site has no reliable public-content boundary. Do not concatenate thousands of pages just to produce a large file. Large output can be expensive to generate, difficult to review, and unpleasant to consume. If documentation changes daily, a stale bundle can be worse than a curated index linking to current pages. A useful test is reversibility: can you regenerate, diff, and remove the file without changing the canonical site? If not, the representation is too entangled with the publication process. ## Why this matters The formats are useful when they make the web easier to inspect. They are harmful when they become a substitute for source quality or a pretext for unsupported AI-search promises. Publish what you can keep current, label the representation accurately, and measure endpoint behavior separately from citations. ## Choosing by site size For a small marketing site, `llms.txt` plus canonical HTML is usually enough. The editorial cost is low, and the file can point to the handful of pages that explain the product. Adding an expanded bundle too early creates more output to keep current without proving that anyone needs it. For a documentation site, page-level twins can be more useful because readers and tools often need one exact reference page. Generate them from the same source as the docs site and make version boundaries explicit. A versioned `/v2/` page should not quietly produce a Markdown twin from `/v1/` content. For a large publication, an expanded file may be operationally expensive. Consider a curated index and stable page routes instead. If you publish a bundle, cap its scope, exclude stale sections, and make its collection date or build version visible. ## A review checklist Before shipping any representation, ask: - Is the source page public and canonical? - Does the Markdown include the answer, not just navigation? - Are source links preserved and visible? - Are private fields excluded before rendering? - Does a changed article invalidate the right cache? - Can the output be diffed in CI? - Have we described it without promising engine behavior? This checklist is deliberately boring. Boring is good when the artifact is consumed outside the browser. ## FAQ ### Is `llms-full.txt` an official standard? It is an implementation pattern associated with the broader llms.txt proposal, not a universal requirement accepted by every AI engine. ### Should `.md` pages have their own canonical URLs? Usually the HTML page remains canonical while the Markdown endpoint is an alternate representation. Document the relationship and avoid creating duplicate indexable pages accidentally. ### Can I block one format with robots.txt? You can express crawler access rules, but matching behavior and support vary. Use server authorization for anything private and test the rules with the relevant vendor guidance. ## Sources - [The /llms.txt file](https://llmstxt.org/) - [Answer.AI llms-txt repository](https://github.com/AnswerDotAI/llms-txt) - [MDN content negotiation](https://developer.mozilla.org/en-US/docs/Web/HTTP/Content_negotiation) - [RFC 9110](https://www.rfc-editor.org/rfc/rfc9110) - [Google JavaScript SEO basics](https://developers.google.com/search/docs/crawling-indexing/javascript/javascript-seo-basics) - [Strand CMS repository](https://github.com/BowTiedSwan/strand) ## FAQ **What is the difference between llms.txt and llms-full.txt?** llms.txt is a concise index of important resources; llms-full.txt is an optional expanded document that may include more content. **Are .md pages better than HTML for AI crawlers?** Not universally. Markdown can be a useful representation, but accessible HTML, clear structure, and documented crawler access remain separate concerns. **Does llms-full.txt replace a website?** No. It can complement canonical pages, but it adds duplication and freshness responsibilities and may be ignored by consumers. **Which format should a small site publish first?** Start with canonical HTML and a concise llms.txt index; add page-level Markdown or an expanded file only when you can keep parity and privacy under control. ## Sources - [The /llms.txt file](https://llmstxt.org/) — llms-txt - [llms-txt reference repository](https://github.com/AnswerDotAI/llms-txt) — Answer.AI - [Content negotiation](https://developer.mozilla.org/en-US/docs/Web/HTTP/Content_negotiation) — MDN - [RFC 9110: HTTP Semantics](https://www.rfc-editor.org/rfc/rfc9110) — IETF - [JavaScript SEO basics](https://developers.google.com/search/docs/crawling-indexing/javascript/javascript-seo-basics) — Google Search Central - [Strand CMS on GitHub](https://github.com/BowTiedSwan/strand) — BowTiedSwan --- # Markdown Twins: Serve a Clean .md of Every Page *By The Strand CMS Team · 2026-08-10 · 9 min read* Canonical: https://www.strandcms.com/blog/markdown-twins > **Summary:** Markdown twins are clean page-level representations served alongside canonical HTML. They can make content easier to inspect, but only if parity, privacy, caching, and canonical-link rules are explicit. Markdown twins are clean `.md` representations of canonical pages. They let a site expose the substance of an article without requiring a consumer to understand the whole application shell. They are a representation pattern, not a universal AI-crawler protocol. This guide belongs to the [llms.txt technical authority](/blog/llms-txt-complete-guide) cluster and focuses on implementation discipline. The shortest safe rule is: generate HTML and Markdown from the same public source, keep HTML canonical, and test that the two versions answer the same question. ## Markdown twins: the useful mental model Think of a page as a resource with more than one representation. HTML is usually the best representation for people and browsers. Markdown can be a cleaner representation for developers, documentation tools, and consumers that need text. [MDN's content-negotiation documentation](https://developer.mozilla.org/en-US/docs/Web/HTTP/Content_negotiation) and [RFC 9110](https://www.rfc-editor.org/rfc/rfc9110) provide the standards vocabulary for representations, selection, and caching. The twin should not become a second editorial source. If a writer changes the answer in HTML but a separate Markdown file keeps the old paragraph, the site now has two conflicting claims. That is a content operations bug, not a formatting detail. ## Why serve a clean Markdown representation A clean text surface can help when: - documentation consumers need predictable text - an agent should inspect the answer without UI noise - a developer wants to diff content in a repository - a site wants a stable machine-readable artifact - a page contains important content behind a client-heavy shell Google's [JavaScript SEO guidance](https://developers.google.com/search/docs/crawling-indexing/javascript/javascript-seo-basics) emphasizes that important content needs to be available in crawlable, renderable output. A Markdown twin can be a useful additional route, but it does not remove the obligation to make the canonical page work. Do not turn this into a ranking claim. A successful `.md` request proves that the endpoint responded. It does not prove that a search or answer engine fetched it, preferred it, or cited it. ## Route design choices There are three common approaches: 1. **Suffix route:** `/guide/intro.md`. Easy to explain and test. 2. **Content negotiation:** the same resource varies by `Accept` header. Semantically elegant, but more complex to cache and debug. 3. **Dedicated text endpoint:** `/api/content/guide/intro`. Useful for an internal consumer, but less obvious as a public alternate representation. Pick one convention and document it. Use stable URLs, return a correct `Content-Type`, and ensure redirects do not silently turn the text route into a browser-only page. If the Markdown URL is public, include the canonical HTML URL in the document itself. ## What to render A good twin normally contains: - title and a concise summary - publication and update dates when public - the article headings and body - tables converted to readable Markdown - source links and related canonical links - a canonical URL near the top Strip or avoid: - account controls and session data - private metadata and internal IDs - draft blocks and preview-only text - repeated navigation and footer boilerplate - analytics parameters in links - hidden content that is not visible to the public reader The output should be useful if copied into a plain text editor. That is a better quality bar than merely removing tags. ## Parity is the real feature Create a parity test that compares the important fields, not just byte output. A Markdown renderer will legitimately differ from HTML in whitespace and syntax. It should not differ in title, primary answer, source list, publication state, or public links. A practical checklist: | Check | Expected result | | --- | --- | | Title | Same visible title | | Summary | Same answer-first summary | | Status | Only published content appears | | Links | Same destinations after tracking removal | | Sources | Same source set and labels | | Dates | Same public publication and update values | | Canonical | HTML URL is explicit | | Privacy | No private or preview fields | Run the test in CI after rendering. Also request the deployed endpoint so caching and routing are part of the check. A local renderer can pass while the edge cache serves yesterday's Markdown. ## Caching and freshness Treat the twin as a cacheable representation of the same content. Use a cache policy that matches the HTML publish process. If HTML updates immediately but Markdown remains stale for hours, a consumer can receive contradictory answers. When using content negotiation, vary caches on the relevant request header as described by HTTP semantics. When using a suffix route, make the relationship explicit in deployment configuration. Log failed renders and unexpected content types. Do not silently fall back to an HTML app shell with a `200` status when a caller requested Markdown. ## Discovery without overclaiming A Markdown twin can be linked from the HTML page, an API response, documentation, or a curated [`llms.txt`](/blog/llms-txt-complete-guide) index. Those links help a consumer discover the route. They do not force consumption. This distinction is worth repeating because “agent-ready” is easy to turn into a slogan. OpenAI and Perplexity publish crawler documentation, but vendor guidance does not establish that every engine uses `.md` routes. Describe what you serve, test what you serve, and avoid claims about behavior you cannot observe. ## Strand's approach Strand stores articles as MDX in Git and generates machine-readable surfaces from the publication source. Its [repository](https://github.com/BowTiedSwan/strand) is the authority for those implementation details. The important design idea is source proximity: when the article, validation, and generated representation share a pipeline, parity is easier to reason about. Read [Introducing Strand CMS](/blog/introducing-strand) for the product-level overview and [llms.txt vs llms-full.txt](/blog/llms-txt-vs-llms-full) for how a page twin differs from a site index or expanded bundle. ## A minimal implementation sketch ```js export function renderMarkdownTwin(post) { if (post.status !== "published") throw new Error("private content"); return [ `# ${post.title}`, `Canonical: ${post.canonicalUrl}`, "", post.summary, "", post.markdownBody, "", "Sources", ... post.sources.map((source) => `- [${source.title}](${source.url})`), ].join("\n"); } ``` This is illustrative, not a complete framework. In production, sanitize links, escape headings, preserve tables, handle missing optional fields, and test that drafts cannot reach the route. The boundary around `post.status` is as important as the Markdown conversion. ## Why this matters Markdown twins make a site easier to inspect when they are boring, predictable, and generated from the same source as the page people read. They do not need a grand theory. They need parity, a clear canonical relationship, sensible caching, and a hard privacy boundary. That is enough to create a useful machine-readable surface without pretending that a file extension controls an AI engine. ## A rollout plan that does not create drift Start with a single public article and render both representations from the same source. Compare the title, summary, headings, links, source list, and updated timestamp. Then publish the route behind a feature flag or a narrow path, request it from outside the development environment, and inspect the actual response headers and body. Once the first route passes, add a small set of representative content: a guide with a table, a post with citations, and a page with an image or code block. These cases reveal renderer assumptions that a plain paragraph will not. Keep the test fixtures public-looking but non-sensitive, and include one unpublished fixture to verify that the route rejects it. Monitor the route after release. Track errors, render failures, cache age, and unexpected status changes. Do not interpret traffic as citation data unless you have a separate, defensible measurement. A Markdown endpoint is a product surface; it deserves the same ownership as an API route. ## Accessibility and reader experience A twin should serve people as well as automated tools. Include a plain link from the canonical page when discovery makes sense, state what the representation is, and keep the text readable without special tooling. Tables need headers, code blocks need language labels where available, and source links should retain their titles. If the Markdown version omits an important visual explanation, link to an accessible textual equivalent or retain the explanation in the canonical page. “Clean” must not mean “less information.” It should mean less interface noise around the same public answer. ## Versioning and content boundaries Documentation teams often publish several versions of the same product. A Markdown twin must preserve the version in its route, title, and canonical link. Otherwise a consumer can fetch a current-looking endpoint that quietly contains older instructions. Put the version in the source model, not only in a navigation label. The same rule applies to audiences. A public guide, a customer-only runbook, and an internal incident note may share a content type but not a distribution policy. Decide whether the twin is generated after authorization and filter before rendering. Never rely on a hidden CSS class or an omitted navigation link to protect restricted text. If the site supports previews, make preview URLs unguessable and require authentication. A public `.md` route should derive only from published records. Add a test that creates a draft with a distinctive sentence and confirms that a request for the public twin cannot find that sentence. ## Observability and failure handling Return an explicit error when a twin cannot be rendered. A `404` for a missing published page is clearer than a `200` app shell. A `500` with a request identifier is clearer than an empty document that looks valid. Log the source revision, renderer version, and route so a stale or malformed representation can be traced. Measure generation failures separately from requests. A low request count does not make a broken route acceptable, and a high request count does not prove citation. These are operational metrics, not search-performance claims. ## Failure modes worth testing Test a missing page, a draft page, a page with a broken source link, and a page whose Markdown renderer fails on a table. Each case should produce an intentional result. If the route silently returns an empty `200`, monitoring will under-report the problem and a consumer may treat the response as complete. Also test deployment boundaries. A CDN may cache the HTML and Markdown routes with different keys. A redirect may remove the `.md` suffix. A proxy may rewrite `Content-Type`. These are ordinary web failures, but they matter more when a machine assumes a stable representation. This is why the endpoint should be treated like a maintained interface. A small, explicit contract is easier to test than a promise that every consumer will infer the same content from every format. ## FAQ ### Should every page have a Markdown twin? No. Start with public pages whose content is stable and useful to text consumers. Expand only after parity and privacy tests are reliable. ### Is a Markdown twin the same as `llms-full.txt`? No. A twin represents one page. An expanded file usually bundles content from many pages into one document. ### Can a Markdown endpoint be indexed separately? It can be, depending on your headers and indexing strategy. Decide deliberately whether it is an alternate representation or a separate public document, and avoid accidental duplicates. ## Sources - [MDN content negotiation](https://developer.mozilla.org/en-US/docs/Web/HTTP/Content_negotiation) - [RFC 9110](https://www.rfc-editor.org/rfc/rfc9110) - [Google JavaScript SEO basics](https://developers.google.com/search/docs/crawling-indexing/javascript/javascript-seo-basics) - [Google SEO Starter Guide](https://developers.google.com/search/docs/fundamentals/seo-starter-guide) - [The /llms.txt file](https://llmstxt.org/) - [Strand CMS repository](https://github.com/BowTiedSwan/strand) ## FAQ **What are Markdown twins?** Markdown twins are clean Markdown representations of canonical web pages, served alongside HTML for readers or tools that need a simpler text surface. **Do Markdown twins improve SEO automatically?** No. They can improve operational accessibility, but they do not automatically improve rankings or guarantee AI citations. **How should a Markdown twin link to HTML?** Include the canonical page URL in the representation and keep the HTML page as the primary public URL unless your indexing design says otherwise. **What should a Markdown endpoint exclude?** Exclude drafts, private fields, account data, noisy navigation, tracking parameters, and anything not intended for public distribution. ## Sources - [Content negotiation](https://developer.mozilla.org/en-US/docs/Web/HTTP/Content_negotiation) — MDN - [RFC 9110: HTTP Semantics](https://www.rfc-editor.org/rfc/rfc9110) — IETF - [JavaScript SEO basics](https://developers.google.com/search/docs/crawling-indexing/javascript/javascript-seo-basics) — Google Search Central - [SEO Starter Guide](https://developers.google.com/search/docs/fundamentals/seo-starter-guide) — Google Search Central - [The /llms.txt file](https://llmstxt.org/) — llms-txt - [Strand CMS on GitHub](https://github.com/BowTiedSwan/strand) — BowTiedSwan --- # AI Crawler robots.txt Guide for 2026 *By The Strand CMS Team · 2026-08-10 · 9 min read* Canonical: https://www.strandcms.com/blog/robots-txt-ai-crawlers > **Summary:** This ai crawler robots.txt guide explains how to use robots.txt as an access policy for documented bot identities. Allowing a crawler never guarantees indexing, ranking, or AI citations. An AI crawler `robots.txt` policy is an access decision, not an AI-search visibility guarantee. This ai crawler robots.txt guide shows how to make that decision without confusing access with citation. `robots.txt` tells compatible crawlers which paths they may request under documented matching rules. It does not authenticate users, remove data from the internet, or force an answer engine to cite a page. This guide is part of the [llms.txt technical authority](/blog/llms-txt-complete-guide) cluster and keeps crawler control separate from machine-readable content formats. The practical workflow is to identify the bot, read its current first-party documentation, choose an allow or disallow policy by content rights, and test the deployed file. ## AI crawler robots.txt: what the file does Google's [robots.txt specification](https://developers.google.com/search/docs/crawling-indexing/robots/robots_txt) describes a text file at the host root with groups of user-agent rules and path directives. The file is a convention honored by compliant crawlers. It is not a firewall and should never be your only protection for confidential material. A basic group looks like this: ```txt User-agent: ExampleBot Disallow: /private/ Allow: /public/ Sitemap: https://example.com/sitemap.xml ``` The exact behavior depends on the crawler, syntax, matching, and deployment. Keep rules narrow. A typo in a broad `Disallow: /` can block more than intended; a typo in a bot name can fail to match at all. ## GPTBot is not every OpenAI crawler OpenAI's [crawler documentation](https://developers.openai.com/api/docs/bots) describes multiple crawler identities and purposes, including GPTBot, OAI-SearchBot, and ChatGPT-User. Do not collapse them into one vague “OpenAI bot” rule without understanding the purpose the vendor assigns to each identity. The rights decision may differ. A publisher could allow a search crawler while disallowing a training crawler, or make the opposite choice based on licensing and business goals. The policy is yours; the vendor documentation tells you which identity you are matching. Allowing GPTBot does not guarantee a page will appear in a ChatGPT answer. It only removes one possible access restriction for that user agent. Retrieval, relevance, indexing, and citation remain separate decisions. ## PerplexityBot and user fetches Perplexity's [crawler guidance](https://docs.perplexity.ai/docs/resources/perplexity-crawlers) documents PerplexityBot and Perplexity-User behavior. Read the current page before writing rules because a search crawler and a user-triggered fetch can have different roles. Do not assume that allowing one Perplexity identity allows all traffic from the service, and do not infer citation from a successful request. A request log can show access; it cannot alone show how a response was composed. ## What about ClaudeBot? Bot names and documentation change. This article deliberately does not present an unsupported ClaudeBot rule as fact. The research note for this batch recorded that the attempted legacy Anthropic URL returned 404, so a current first-party source must be retrieved before making a policy claim about that identity. That is the correct editorial answer: if the source is unavailable, mark the claim unresolved or cut it. A plausible bot name is not evidence. ## Rules versus security Use robots.txt for crawl preferences and access signaling. Use authentication, authorization middleware, signed URLs, and storage permissions for private content. A disallow directive does not stop a malicious client from requesting a URL, and it does not erase a URL that has already been copied elsewhere. The same boundary applies to `llms.txt`: the proposed file can describe public resources, but it is not an access-control system. See [The Complete llms.txt Guide](/blog/llms-txt-complete-guide) for the distinction. ## A reversible policy process 1. Inventory public, licensed, sensitive, and private paths. 2. List the vendor crawler identities that matter to your business. 3. Read each vendor's current first-party documentation. 4. Write the smallest rule group that matches your decision. 5. Deploy to the root and request it from a clean network. 6. Validate syntax and check server logs for unexpected matches. 7. Revisit the policy when rights, products, or vendor identities change. Keep the file in version control. Require review for changes to `Disallow: /`, wildcard rules, or sensitive path groups. Add a deployment check that confirms the production file is the one you reviewed. ## Rendering comes after access A crawler that is allowed to fetch a page can still receive a weak representation. Google's [JavaScript SEO guide](https://developers.google.com/search/docs/crawling-indexing/javascript/javascript-seo-basics) explains why rendering and crawlable output matter. If the answer only appears after a fragile client-side request, changing robots rules will not fix the page. For a cleaner representation strategy, see [Markdown Twins](/blog/markdown-twins). The format can complement HTML, but it does not turn an inaccessible or low-quality page into a trustworthy source. ## Strand's implementation notes Strand keeps the site configuration and content workflow in a repository, which makes robots changes reviewable alongside the publication. Its [GitHub repository](https://github.com/BowTiedSwan/strand) is the source for Strand-specific claims. The broader lesson is operational: a rule that can be diffed, validated, and reverted is easier to trust than a setting hidden in an unreviewed dashboard. Read [Introducing Strand CMS](/blog/introducing-strand) for the product overview. ## Why this matters Crawler policy is one of the few AI-search controls a publisher can make explicit, but it is easy to oversell. Name the bot, state the purpose, keep private content behind real authorization, and document what an allow or block decision means. A clean policy reduces accidental access; it does not promise visibility. ## Testing a production policy Do not validate only the file in a local checkout. Request the production URL, inspect redirects, and record the exact response. Check the root location, because a file under `/docs/robots.txt` does not control the host root. Then exercise representative paths from each rule group and inspect logs for the user-agent strings you care about. Keep a small test matrix: | User agent | Public article | Private area | Expected | | --- | --- | --- | --- | | Search crawler | allowed | blocked | public content remains discoverable | | Search crawler | blocked | blocked | restricted section stays out | | Vendor search bot | policy decision | policy decision | matches rights choice | | Unknown client | server authorization | server authorization | robots is not security | A test matrix cannot simulate every crawler, but it catches the dangerous mistakes: a global block, a missing root file, and the assumption that a user-agent string is an authenticated identity. Treat it as a deployment guard, not a proof of vendor behavior. ## Keep the policy legible Comments can explain why a rule exists, but they are not a substitute for a change record. Store the policy with the site, link the relevant vendor documentation in the review note, and assign an owner for rights decisions. When a vendor changes crawler names or purposes, update the rule deliberately rather than adding every newly mentioned string. The cleanest policy is one a future editor can understand in a minute. Broad, unexplained blocks create operational debt; broad allows create rights risk. Narrow groups and explicit rationale make both conversations easier. ## Policy examples by objective A publisher that wants broad public discovery might allow a documented search crawler on article paths while disallowing administrative routes. A company with contractual restrictions might disallow a training-related identity while continuing to serve public pages to ordinary search crawlers. A private application might disallow all automated paths, but it still needs authentication because robots rules are advisory. These examples are decision shapes, not recommendations for every site. The right choice depends on the rights you have and the behavior you want. Write the rationale next to the rule so a future editor can distinguish an intentional block from a forgotten experiment. ## User-agent matching pitfalls User-agent matching is often more subtle than the visual size of the file suggests. A group for one token may not cover a related token. Wildcards and path matching need to be tested against the crawler's documented parser. A rule that works in a local validator can still be deployed to the wrong host, wrong protocol, or wrong environment. Check both apex and `www` hosts if they serve different responses. Check redirects because a crawler may fetch the first host's policy before following to another. Check staging and production independently, and make sure a deployment does not replace the production file with a default generated file. ## Keeping crawler policy and content policy together Robots rules should be reviewed with the content model. If a new route contains customer data, the secure boundary belongs in application authorization first; the robots change is only a supplementary signal. If a new public article is intended for discovery, verify that its canonical HTML, sitemap, internal links, and machine-readable alternatives agree. The [llms.txt guide](/blog/llms-txt-complete-guide) covers the index convention. The [Markdown Twins](/blog/markdown-twins) guide covers alternate representations. Treat the three artifacts as related but independent: one maps resources, one represents content, and one expresses crawler preferences. ## A change-review template When changing the file, record the date, the user-agent group, the paths affected, and the reason. Link the vendor page that supports any claim about purpose or matching. State whether the change affects search crawling, training access, user-triggered fetches, or only a private application path. This prevents a short configuration diff from hiding a large policy decision. After deployment, request the root file from each host and compare it with the reviewed version. Test one allowed and one disallowed path where your tooling supports it. If the result differs, stop the rollout and inspect the host, redirect, cache, and parser assumptions before changing more rules. Document the fallback when a rule cannot be evaluated. The answer should be the site's secure server behavior, not an assumption that the crawler will obey a text file. Clear failure handling keeps robots policy from becoming a false sense of protection. ## Incident response for an accidental block If a deployment blocks an important crawler, first restore the last known-good file rather than layering on more exceptions. Confirm the production host, purge the relevant cache, and request the file again. Then check whether the affected page is still accessible to ordinary readers and whether the block came from robots, authentication, a firewall, or a rendering failure. Record the cause and add a regression check before making a permanent policy change. An accidental allow deserves the same care. Identify which paths were exposed, rotate secrets if private material was reachable, and fix authorization at the application boundary. Changing robots can reduce future requests, but it cannot recall a response already sent to a client. ## FAQ ### Should I allow every AI crawler? No. Make the decision based on content rights, business goals, and the vendor's documented purposes. There is no obligation to use a blanket allow rule. ### Does robots.txt affect Google rankings directly? It can affect whether compliant crawlers can request content, but it is not a ranking lever by itself. Make sure important public pages are accessible and useful. ### How often should I review the file? Review it when vendor documentation, content rights, site routes, or deployment architecture changes. A periodic check is sensible for active publications. ## Sources - [Google robots.txt specification](https://developers.google.com/search/docs/crawling-indexing/robots/robots_txt) - [OpenAI crawler documentation](https://developers.openai.com/api/docs/bots) - [Perplexity crawler documentation](https://docs.perplexity.ai/docs/resources/perplexity-crawlers) - [Google JavaScript SEO basics](https://developers.google.com/search/docs/crawling-indexing/javascript/javascript-seo-basics) - [Google SEO Starter Guide](https://developers.google.com/search/docs/fundamentals/seo-starter-guide) - [Strand CMS repository](https://github.com/BowTiedSwan/strand) ## FAQ **What is the right robots.txt rule for AI crawlers?** There is no universal rule. Decide by bot purpose, content rights, and business need, then test the exact user-agent matching behavior. **Does allowing GPTBot guarantee ChatGPT citations?** No. OpenAI documents separate crawler purposes, and access permission does not guarantee retrieval, inclusion, or citation. **Is ClaudeBot documented here?** This guide does not assert a ClaudeBot policy without a live first-party source. Check current vendor documentation before adding a rule. **Can robots.txt protect private content?** No. Use authentication and server authorization for private content; robots.txt is not a security boundary. ## Sources - [robots.txt Specifications](https://developers.google.com/search/docs/crawling-indexing/robots/robots_txt) — Google Search Central - [Overview of OpenAI Crawlers](https://developers.openai.com/api/docs/bots) — OpenAI - [Perplexity Crawlers](https://docs.perplexity.ai/docs/resources/perplexity-crawlers) — Perplexity - [JavaScript SEO basics](https://developers.google.com/search/docs/crawling-indexing/javascript/javascript-seo-basics) — Google Search Central - [SEO Starter Guide](https://developers.google.com/search/docs/fundamentals/seo-starter-guide) — Google Search Central - [Strand CMS on GitHub](https://github.com/BowTiedSwan/strand) — BowTiedSwan --- # AI Crawlers JavaScript: Why SPAs Lose Citations in 2026 *By The Strand CMS Team · 2026-07-20 · 7 min read* Canonical: https://www.strandcms.com/blog/ai-crawlers-javascript-spa > **Summary:** If the important copy only appears after client-side rendering, many crawlers will miss it or see it late. Ship the answer in HTML first. If the main article text only appears after JavaScript runs, you are asking crawlers to do extra work for no gain. That is a bad trade on a news page, a blog post, or any page you want cited in an answer engine. This is the core of the AI crawlers JavaScript problem. This piece sits in the [best CMS for AI search](/blog/best-cms-for-ai-search-visibility) cluster. The fix is not mysterious. Put the content in the HTML first. ## What breaks on JavaScript-heavy sites Google documents the limits of JavaScript rendering. In plain English, that means the crawler may not see your full page at the same moment a browser user does. OpenAI and Perplexity both publish crawler docs, which is a useful reminder that not every bot behaves the same way. If the headline is visible but the body is assembled later, you are already taking a risk. ## What to do instead - Server render the article body. - Use static generation when the page does not need live data. - Keep the first HTML response useful on its own. - Use JavaScript for interactions, not for hiding the actual article. That is why Strand keeps the published surface static and writes the source as MDX in Git. The agent can still help with the draft. The crawler should not need a tour guide. ## Common SPA failure modes The most common problem is not a total failure. It is partial visibility. The headline appears, but the body text does not. The article body appears, but the sources load later. The article renders, but the page shell makes it hard to tell where the content starts. The metadata changes on the client after load, which creates a gap between what the page says and what the crawler sees. None of those bugs sounds dramatic on its own. Together, they make citation harder. ## AI crawlers JavaScript: the safest default If your page needs to be cited, assume the crawler should be able to read the important text before any client-side code runs. That does not mean banning JavaScript. It means using it for interactions, not for hiding the article body, summary, or sources. A clean server response is easier to test, easier to compare against the live page, and easier for an answer engine to quote without missing the point. The most common problem is not a total failure. It is partial visibility. The headline appears, but the body text does not. The article body appears, but the sources load later. The article renders, but the page shell makes it hard to tell where the content starts. The metadata changes on the client after load, which creates a gap between what the page says and what the crawler sees. None of those bugs sounds dramatic on its own. Together, they make citation harder. ## A migration path that does not require a rewrite You do not need to scrap the whole frontend to fix this. Start by rendering the article body on the server. Keep interactive widgets in JavaScript. Move the canonical summary and sources into the HTML response. Leave the app shell for things that need it. That usually gets you most of the benefit without forcing a redesign. ## Why this matters in practice The more predictable the HTML, the easier it is for a model to lift the right answer. That matters for pages that want to be cited, but it also matters for team sanity. A static source of truth is easier to debug than a client-only stack that changes shape after hydration. ## A quick test Ask yourself a simple question: if JavaScript failed, would the page still tell the reader anything useful? If the answer is no, the page needs work. You can also test the raw HTML with `curl` or inspect the page source. If the core article body is missing there, a crawler may have the same problem. ## When JavaScript is fine JavaScript is not the problem by itself. Navigation, search, forms, tabs, and comments can all be fine. The problem starts when JavaScript is the only place the article text exists. That distinction matters. Most teams do not need to remove JavaScript. They need to stop using it as the container for the content itself. ## Why this matters A crawler that sees the right HTML on the first pass can quote the page without guessing. That makes the result more stable, which matters more than people think. A stable page is easier to index, easier to test, and easier to share with a teammate who is trying to understand what went wrong. The practical lesson is simple. Keep the article body, the summary, and the sources in the initial response. Use JavaScript for menus, tabs, toggles, and other interaction. That is usually enough to keep the page friendly for both readers and answer engines. If you want the version of this argument tied to a publishing system, read [Introducing Strand CMS](/blog/introducing-strand). ## A practical rollout If you are fixing a site that was built as a client-heavy SPA, do not try to redesign it in one sprint. Start by rendering the content that matters most. Make the article body available in the HTML response. Move the canonical summary there too. Then do the same for the sources and author block. Once that works, test with scripts disabled. If the page still communicates the main point, you are close. If it turns into a blank shell, the job is not done yet. The last step is to keep the interactive pieces where they belong. Search widgets, tabs, and modals can stay. The article itself should not depend on them. ## Why this matters in practice A lot of teams treat JavaScript rendering as a frontend preference. For AI visibility, it is also a publishing decision. The page that ships in HTML is the one a crawler can trust first. That is why boring rendering choices often beat clever ones. The cleaner the source, the less the model has to guess. ## A quick diagnosis If you are not sure whether a page is too dependent on JavaScript, check the source in a plain browser view. If the article body is missing, the page is too dependent. If the article body is present but the sources are not, the page is still risky. If the article body is present and readable, you are already in much better shape. ## A safer build path The cleanest fix is usually boring: render the article on the server, keep the interactive bits in JavaScript, and avoid changing the meaning of the page after load. That gives you a page that is easier for crawlers to read and easier for humans to share. It also makes debugging less annoying, which is reason enough on its own. ## When to stop worrying You do not need to ban JavaScript from the site. You need to stop using it as the only way to deliver the content. If the headline, the answer, the summary, and the sources all exist in the HTML response, most of the risk is already gone. If the page still feels like a normal article when scripts are blocked, that is usually a good sign. If you want the version of this argument tied to a publishing system, read [Introducing Strand CMS](/blog/introducing-strand). ## FAQ ### Is SSR always required? No, but it is the safest default for content that needs to be crawled. ### Do crawlers run JavaScript perfectly now? Not perfectly, and not consistently enough to treat client-only rendering as safe for important content. ### Can I keep my SPA and still be search-friendly? Yes, if the article body is present in the initial HTML and the app shell does not hide it. ## Sources - [Understand JavaScript SEO Basics](https://developers.google.com/search/docs/crawling-indexing/javascript/javascript-seo-basics) - [How Google Interprets the robots.txt Specification](https://developers.google.com/search/docs/crawling-indexing/robots/robots_txt) - [Overview of OpenAI Crawlers](https://developers.openai.com/api/docs/bots) - [Perplexity Crawlers](https://docs.perplexity.ai/docs/resources/perplexity-crawlers) - [GitHub - BowTiedSwan/strand](https://github.com/BowTiedSwan/strand) - [Introducing Strand CMS: publishing built for agents](https://strandcms.com/blog/introducing-strand) ## FAQ **Are JavaScript sites always invisible to crawlers?** No. But they are riskier if the main content depends on client-side rendering or delayed hydration. **What is the safest fix for a SPA?** Server render or statically render the content that you want crawlers to read. **Do I need to remove JavaScript entirely?** No. You just should not hide the article behind JavaScript when the article itself is the thing you want indexed. ## Sources - [Understand JavaScript SEO Basics](https://developers.google.com/search/docs/crawling-indexing/javascript/javascript-seo-basics) — Google - [How Google Interprets the robots.txt Specification](https://developers.google.com/search/docs/crawling-indexing/robots/robots_txt) — Google - [Overview of OpenAI Crawlers](https://developers.openai.com/api/docs/bots) — OpenAI - [Perplexity Crawlers](https://docs.perplexity.ai/docs/resources/perplexity-crawlers) — Perplexity - [GitHub - BowTiedSwan/strand](https://github.com/BowTiedSwan/strand) — BowTiedSwan - [Introducing Strand CMS: publishing built for agents](https://strandcms.com/blog/introducing-strand) — Strand CMS --- # Best CMS for AI Search: 7 Options Ranked for 2026 *By The Strand CMS Team · 2026-07-20 · 12 min read* Canonical: https://www.strandcms.com/blog/best-cms-for-ai-search-visibility > **Summary:** If AI search visibility matters, compare CMSs on crawlable HTML, structured content, file or Git workflows, and whether they make citations easy. If the question is "best CMS for AI search" the answer is not the one with the prettiest dashboard. It is the one that gives crawlers clean HTML, gives editors a structured model, and gives you a source format that does not turn into a mess the second the page changes. This piece sits in the same cluster as [Introducing Strand CMS](/blog/introducing-strand). This roundup compares six options on that basis. It is not a popularity contest. It is a look at what the system actually emits, how the content is stored, and how much work it takes to keep the page legible to answer engines. ## Best CMS for AI Search: quick verdict | Tool | Best for | Pricing | Open source? | | --- | --- | --- | --- | | Strand | Git-native publishing, clean Markdown twins, AI search artifacts | Open source, self hosted | Yes | | Ghost | Newsletters, memberships, and a polished hosted publishing flow | Paid hosted plans available; self hosting also possible | Yes | | WordPress | General publishing with a huge ecosystem | Free core; hosting and plugins vary | Yes | | Sanity | Structured content models and custom frontends | Free tier plus paid usage | No | | Strapi | API-first teams that want a strong admin panel and extensibility | Free self-hosted core plus paid cloud | Yes | | Payload | TypeScript teams that want CMS and app framework in one codebase | Free core plus paid hosting and services | Yes | ## 6-factor verdict ## What the vendor docs actually say The comparison only stays honest if you look at the current docs, not old marketing pages. Ghost's docs center publishing and memberships. WordPress documentation centers general publishing and its ecosystem. Sanity documents structured content and custom frontends. Strapi documents content modeling, APIs, preview, and deployment. Payload describes itself as a docs-first, developer-centric CMS with a built-in admin panel and app-framework feel. That matters because AI search visibility is mostly about the output surface. If the docs emphasize structured content or predictable rendering, the stack usually gives you a cleaner path to crawlable pages. If the docs emphasize plugins or app composition, you need stricter editorial and rendering discipline to keep the output stable. ## The filter I would use A CMS wins this comparison when it helps you do three things at once: keep the HTML crawlable, keep the content model explicit, and keep the source of truth stable. If a tool does only one of those well, you still have work to do. The honest tradeoff is that Ghost and WordPress win on familiar publishing ergonomics, while Sanity, Strapi, and Payload win when the site is really an app with a content layer. Strand wins when you care most about Git-native source control, clean Markdown twins, and a machine-readable output surface that stays close to the draft. A final sanity check is to ask where the stack creates drift. If the answer is preview, metadata, or rendering, you need more discipline. If the answer is the source format itself, the stack is probably a better fit for AI search work. That is the difference between a stack that merely stores content and a stack that keeps the publication readable to machines. | Tool | Crawlable HTML | Structured content | File or Git workflow | AI-search extras | Operational drag | Standout win | | --- | --- | --- | --- | --- | --- | --- | | Strand | Strong | Strong | Strong | Strong | Low | Best Git-native stack | | Ghost | Strong | Good | Medium | Medium | Low | Best for newsletters | | WordPress | Strong | Medium | Medium | Medium | Medium | Biggest ecosystem | | Sanity | Good | Strong | Medium | Medium | Medium | Best content modeling | | Strapi | Good | Strong | Medium | Medium | Medium | Best API-first admin | | Payload | Good | Strong | Medium | Medium | Medium | Best app-framework fit | ### Strand Strand is the cleanest fit if your goal is to make content easy for agents and crawlers to ingest. Articles live as MDX in Git, the site emits a clean Markdown version of every page, and the platform ships `llms.txt` plus a validation-first content schema. That is a simple stack to reason about, and it keeps the source of truth where humans already know how to review it. If you want the narrowest path from draft to published page, start with [Introducing Strand CMS](/blog/introducing-strand) and the [GitHub repo](https://github.com/BowTiedSwan/strand). ### Ghost Ghost is the best pick if your publication depends on newsletters and memberships. Its own homepage makes that the center of the product, not a side feature. It is open source, has a strong hosted offering, and is built for people who want to ship a professional publication without assembling a stack from scratch. Choose Ghost if your main job is publishing and monetization, not experimenting with the content model. ### WordPress WordPress still wins on raw ecosystem depth. If you need a plugin for a niche workflow, someone has probably built one already. That can save time when you need a straightforward site and do not want to invent every layer yourself. The tradeoff is that the ecosystem is broad enough to make the publishing surface messy if you are not strict about templates, plugins, and rendering rules. ### Sanity Sanity is strongest when your content model is the product. Its docs frame it as a content operating system, and the platform is built around structured content, custom frontends, and AI-aware tooling. If your team cares about schemas and content reuse more than the default editor experience, Sanity is hard to beat. ### Strapi Strapi is a good fit when the admin panel matters and the API layer has to be flexible. The docs show a strong CMS plus REST and GraphQL support, a visual content-type builder, and a lot of extension points. It is a practical choice for teams that want a real CMS but still expect to wire it into a custom app. ### Payload Payload is the best fit for TypeScript-heavy teams that want the CMS inside the app instead of beside it. The docs describe it as a Next.js fullstack framework with a built-in admin panel, REST and GraphQL APIs, auth, access control, uploads, and preview. That is attractive if your team already thinks in code first. ## The decision in plain English If you are a publisher and newsletters are the product, Ghost is still the easiest fit. If you are building a custom app and need the CMS to live inside the codebase, Payload or Sanity may make more sense. If you want the shortest path from source text to a clean, machine-readable publication, Strand is the least annoying option in the group. If your team values raw ecosystem size over neatness, WordPress still belongs in the conversation. ## Related links that matter - [Best CMS for AI Search](/blog/best-cms-for-ai-search-visibility) - [Introducing Strand CMS](/blog/introducing-strand) - [Get Cited by ChatGPT](/blog/get-cited-by-chatgpt-perplexity) - [GEO, Explained](/blog/geo-explained) - [Does Your Site Have /llms.txt? How to Check and Fix It in 10 Minutes](/blog/does-your-site-have-llms-txt) - [Why AI Crawlers Can't Read Your JavaScript Site in 2026](/blog/ai-crawlers-javascript-spa) - [Structured Data AI Search: Speakable, FAQPage, and E-E-A-T](/blog/structured-data-for-ai-citations) ## What I would choose If AI search visibility is the main requirement, Strand is the shortest path because it minimizes transformation between source text and published page. If you need a media business with subscriptions, Ghost is the practical winner. If you want a broad ecosystem and a familiar stack, WordPress is still the default answer. If your team cares more about app architecture than publishing mechanics, Sanity, Strapi, and Payload each make sense in different ways. ## How I would think about the tradeoffs A lot of CMS comparisons collapse into feature checklists that do not matter once the site goes live. A better test is how much work the stack creates after the first publish. A good stack should make it hard to publish broken metadata. It should make it easy to see the source of truth. It should not require three separate systems to keep the article, preview, and published page aligned. That is where the Git-native option is attractive. It reduces the chance that the content model and the actual page drift apart. ## Which stack fits which team If you are a solo founder or small editorial team, the main thing you want is a system that does not create extra work every week. If you are running a publication, the main thing you want is a clean handoff from draft to publish and a monetization path that does not fight the editorial process. If you are shipping a product site, the CMS should not force your app team into weird workarounds just to make a blog post live. Those are basic needs, but they are the ones that decide whether the stack feels calm or annoying six months later. ### In practice Strand is strongest when the team wants the published page to stay close to the source text. Ghost is strongest when the site is a publication first and a software product second. WordPress is strongest when the team wants a familiar default and a huge plugin market. Sanity is strongest when the content model matters more than the default editor. Strapi and Payload make the most sense when the team already expects to build around code. That is the useful way to think about the comparison. Pick the stack that creates the fewest excuses between the draft and the page. ## How I would decide If the team ships mostly written content, I would bias toward the stack that keeps editorial control simple and the output stable. If the site is a publication with a revenue model, I would favor the one that makes the publishing flow calm and predictable. If the site is a product with a content layer, I would pick the one that fits the app without forcing the app into weird shapes. The point is not to worship one stack. It is to reduce the number of places where the page can drift away from the source text. ## What each stack gives up Strand gives up the broadest plugin ecosystem, but it buys back a cleaner source of truth. Ghost gives up some flexibility, but it buys back a publishing flow that feels made for writing. WordPress gives up some structure, but it buys back familiarity and a massive ecosystem. Sanity gives up the simplicity of a single built-in publishing opinion, but it buys back a lot of control over content modeling. Strapi and Payload give up some out-of-the-box calm, but they buy back API flexibility and room to build around the app. That trade is the real one. Every stack has a bill. The question is which bill you are willing to pay. ## The shortest honest answer If you want the fewest moving parts for AI search work, Strand is the cleanest fit because it keeps the source and the published surface close together. If your business depends on newsletters or subscriptions, Ghost is still a strong choice. If your team values ecosystem size over neatness, WordPress is the default answer. If you are building an app and the content layer is just one part of the product, Sanity, Strapi, or Payload may fit better. The right answer is the one that lets you publish good pages without making the editorial process miserable. ## A practical decision matrix A small editorial team usually cares most about low friction and clear source control. A publication usually cares most about reliable publishing and a monetization path. A product team usually cares most about fitting the CMS into the codebase without turning every content update into an engineering request. That is why the best stack can change depending on the team. The important thing is to notice where the work lives. If the work lives in the editor, pick the stack that makes editorial work smoother. If the work lives in the app, pick the stack that keeps the app flexible. ## Related links that matter - [Best CMS for AI Search](/blog/best-cms-for-ai-search-visibility) - [Introducing Strand CMS](/blog/introducing-strand) - [Get Cited by ChatGPT](/blog/get-cited-by-chatgpt-perplexity) - [GEO, Explained](/blog/geo-explained) - [Does Your Site Have /llms.txt? How to Check and Fix It in 10 Minutes](/blog/does-your-site-have-llms-txt) - [Why AI Crawlers Can't Read Your JavaScript Site in 2026](/blog/ai-crawlers-javascript-spa) - [Structured Data AI Search: Speakable, FAQPage, and E-E-A-T](/blog/structured-data-for-ai-citations) ## Why this matters AI search systems do not reward vague strategy decks. They reward pages that are easy to fetch, easy to parse, and easy to quote. Google documents the rendering limits of JavaScript-heavy sites. OpenAI and Perplexity both publish crawler guidance. Google also says `llms.txt` does not change Search visibility, which is exactly why the publishing format matters more than the slogan. That is the real question this roundup answers: how much of your content model survives contact with a crawler. ## FAQ ### Is a headless CMS automatically better for AI search? No. Headless only helps if the output is still crawlable and the content model stays clean. A messy headless setup can be worse than a simple static site. ### Is Ghost bad for AI search? No. Ghost is strong if your site is mostly publishing and newsletters. It just solves a different problem than a Git-native system. ### Why rank Strand first? Because this comparison is about AI search visibility, and Strand removes the most friction between source text, rendered HTML, and machine-readable output. ## Sources - [GitHub - BowTiedSwan/strand](https://github.com/BowTiedSwan/strand) - [Getting Started With Ghost - Ghost Developer Docs](https://ghost.org/docs/) - [Documentation - WordPress.org](https://wordpress.org/documentation/) - [Home | Sanity Docs](https://www.sanity.io/docs/) - [Strapi 5 Docs | Strapi 5 Documentation](https://docs.strapi.io/) - [What is Payload? | Documentation | Payload](https://payloadcms.com/docs) - [Google's Guide to Optimizing for Generative AI Features on Google Search](https://developers.google.com/search/docs/fundamentals/ai-optimization-guide) ## FAQ **What makes a CMS good for AI search visibility?** It should render content in crawlable HTML, keep the structure clean, and make it easy to publish stable source files or Markdown. **Does llms.txt replace normal SEO?** No. Google says llms.txt does not affect Search visibility, so treat it as a helper file, not a ranking shortcut. **Which CMS is best if I want Git-native publishing?** Strand is the clearest fit because the content lives as MDX in Git and ships a clean Markdown surface for agents and crawlers. ## Sources - [GitHub - BowTiedSwan/strand](https://github.com/BowTiedSwan/strand) — BowTiedSwan - [Getting Started With Ghost - Ghost Developer Docs](https://ghost.org/docs/) — Ghost - [Documentation - WordPress.org](https://wordpress.org/documentation/) — WordPress - [Home | Sanity Docs](https://www.sanity.io/docs/) — Sanity - [Strapi 5 Docs | Strapi 5 Documentation](https://docs.strapi.io/) — Strapi - [What is Payload? | Documentation | Payload](https://payloadcms.com/docs) — Payload - [Google's Guide to Optimizing for Generative AI Features on Google Search](https://developers.google.com/search/docs/fundamentals/ai-optimization-guide) — Google --- # LLMs.txt Setup: Check and Fix /llms.txt Fast *By The Strand CMS Team · 2026-07-20 · 9 min read* Canonical: https://www.strandcms.com/blog/does-your-site-have-llms-txt > **Summary:** LLMs.txt setup is simple: point machines at the pages that explain the site, keep the file short, and treat it as a helper, not a ranking trick. A good llms.txt setup starts with a small file and a clear reason for each link. The fastest way to check for `llms.txt` is to open `https://yourdomain.com/llms.txt` or run `curl -s https://yourdomain.com/llms.txt`. If the file exists, you should see a short index of pages, not a wall of marketing copy. This belongs in the same cluster as [best CMS for AI search](/blog/best-cms-for-ai-search-visibility). If you have to squint to understand what the file is for, it is probably too long. ## LLMs.txt setup: what it is for The proposal behind `llms.txt` is simple. Give LLMs a small, readable entry point to your site so they do not have to infer the shape of the whole thing from a giant homepage. The file is supposed to help at inference time, not replace the public page itself. It is also not a ranking hack. Google says `llms.txt` does not affect Search visibility. ## What to check 1. Does `/llms.txt` return a 200 response? 2. Does it list the pages you actually want machines to read? 3. Are those pages canonical and up to date? 4. Does `robots.txt` still allow crawlers to reach them? 5. Does the file stay short enough to be useful? If the answer to any of those is no, fix that first. ## What to put in it A useful `llms.txt` file is not long. It usually includes: - a short site description - the main docs or blog index - one or two key product pages - any pages that are especially useful to answer engines The point is guidance, not exhaustiveness. ## A few concrete examples A news site might list the homepage, the latest stories, and the editorial standards page. A documentation site might list the docs home, the API reference, and the getting-started guide. A product site might list the homepage, pricing, integrations, and one strong use case page. The common pattern is the same. Point the model toward the pages that best explain the site and let the rest of the site stay discoverable through normal crawling. ## What not to do Do not paste the entire site map into `llms.txt`. That turns a helpful index into another file nobody wants to maintain. Do not use the file as a substitute for crawlable HTML. If the page itself is broken, the helper file will not rescue it. Do not treat the file as proof that your SEO work is done. Google has already said it does not change Search visibility. ## Where Strand fits Strand keeps this simple by pairing a Git-native content model with a machine-readable surface. That is useful because the human process and the crawler surface stay close together. If the site already has a source-of-truth article like [Introducing Strand CMS](/blog/introducing-strand), `llms.txt` can point there instead of forcing a model to infer the structure from scratch. ## A quick fix path If your site does not have one yet, start with the pages that matter most: - the homepage - the main docs or blog hub - one or two high-value articles - a concise explanation of what the site is for Then keep the list small and boring. That is better than trying to make the file impressive. ## A complete rollout A good rollout is boring on purpose. Start with the pages that already do the best job explaining the site, then make sure the file points there and nowhere else. If you run a documentation site, that usually means the docs home and the getting-started page. If you run a product site, that usually means the homepage, pricing, and one or two pages that answer the most common questions. Once the file is live, treat it like a small control surface. If the homepage changes direction, update the description. If the docs get reorganized, update the links. If a page stops being useful, remove it. The file should reflect the site as it is now, not the version you wish it still was. ## A practical file shape A useful `llms.txt` file often has four parts: a one-sentence description of the site, a docs or blog hub, a small list of high-value pages, and links that help a model understand what to read next. That format keeps the file short and easier to maintain than a full sitemap dump. If your site changes often, build the file around pages that are stable and explanatory. A homepage, a docs index, a pricing page, and a few cornerstone posts usually do more for the model than twenty near-duplicates. ## How to audit it Open the file, read it as if you were an external assistant, and ask whether the path from the file to the page is obvious. Then check whether `robots.txt` still allows access and whether the canonical page is still the best entry point. If a page became stale, remove it rather than hoping the file will compensate. ## What to list first The pages that usually earn their place are the homepage, the docs or blog hub, a pricing page, and the single best explainer for what the site is. If the site is a publication, the editorial policy page is often more useful than a giant archive. If the site is a product, the integration or use-case page is often more useful than the marketing homepage alone. ## How to keep it honest Update the file whenever the site structure changes. If the page no longer reflects the current site, remove it. The value of `llms.txt` is that it stays short enough to inspect quickly. The moment it starts looking like a sitemap, it stops being useful. A good rollout is boring on purpose. Start with the pages that already do the best job explaining the site, then make sure the file points there and nowhere else. If you run a documentation site, that usually means the docs home and the getting-started page. If you run a product site, that usually means the homepage, pricing, and one or two pages that answer the most common questions. Once the file is live, treat it like a small control surface. If the homepage changes direction, update the description. If the docs get reorganized, update the links. If a page stops being useful, remove it. The file should reflect the site as it is now, not the version you wish it still was. ## What good looks like A good `llms.txt` file is short enough to scan in a few seconds and clear enough that a model can use it without guessing. It should not feel like a project in itself. It should feel like a map you could hand to someone and trust them to get moving. That is why Strand ships this kind of surface next to the content model. The source of truth and the helper file stay close together, so the maintenance cost stays low. ## A practical maintenance loop Check the file whenever the homepage changes, the docs tree gets reshaped, or the pages that matter most change their purpose. If a link stops earning its place, remove it. If a better source-of-truth page appears, swap it in. That sounds trivial. It is. But that trivial work is what keeps the file useful instead of decorative. ## A complete example A small software site might list the homepage, a getting-started guide, the main docs index, and the pricing page. A content site might list the homepage, the article hub, the editorial policy page, and one or two cornerstone posts. Those are the pages a model is most likely to need. The file should guide the model to the pages that explain the site in the fewest steps. It should not compete with the sitemap or duplicate every internal link on the site. ## Keep the list short A company home page, a docs hub, a pricing page, and one or two high-value product pages are often enough. If the file starts to look like an internal sitemap, it has gone too far. The better test is simple. If a new visitor or a model can use the file to find the pages that explain the site fastest, the file is doing its job. If it just mirrors the whole site, it is wasting space. ## Why brevity matters A useful `llms.txt` file makes the site easier to map without pretending to be something it is not. That matters because a model only gets a small budget of attention. If you spend that budget on a noisy index, you waste the part that should have pointed the model at your best pages. A short file also makes maintenance easier. You can tell at a glance whether the homepage changed, whether the docs index is stale, and whether a product page still belongs in the list. That is the kind of boring maintenance work that keeps the file honest over time. If the file starts to feel like a sitemap clone, trim it. The goal is to guide, not to impress. If you later introduce `llms-full.txt`, treat it as a separate surface with a separate maintenance cost. Start with the small index first, because the short file is the one that stays honest. The best test is whether a person unfamiliar with the site could use the file to find the most useful page in one hop. If not, the list is still too broad. A file that stays useful usually names the same few pages for the same reason every time: the site’s purpose, the most useful hub, and the page a new reader should open next. That consistency matters more than clever wording because the assistant reading the file needs a stable map, not a creative brief. If you are unsure whether a link belongs, ask whether it would still deserve a place if the site had to be explained in one minute. If the answer is yes, keep it. If the answer is only that the page exists, drop it. ## What crawler docs actually say OpenAI and Perplexity both publish crawler docs. Google publishes robots and AI guidance. The message across all of them is consistent: public pages need to be easy to fetch, easy to parse, and hard to misunderstand. `llms.txt` is one small way to help, but it only works if the pages behind it are already solid. ## FAQ ### Is llms.txt required? No. It is optional. ### Does Google use llms.txt for ranking? No. Google says it does not affect Search visibility. ### Should I add llms-full.txt too? Only if you have a reason. Start with the smaller index first. ## Sources - [The /llms.txt file](https://llmstxt.org/) - [Latest Google Search Documentation Updates](https://developers.google.com/search/updates) - [Overview of OpenAI Crawlers](https://developers.openai.com/api/docs/bots) - [Perplexity Crawlers](https://docs.perplexity.ai/docs/resources/perplexity-crawlers) - [How Google Interprets the robots.txt Specification](https://developers.google.com/search/docs/crawling-indexing/robots/robots_txt) - [Strand CMS on GitHub](https://github.com/BowTiedSwan/strand) ## FAQ **What is llms.txt?** It is a proposed `/llms.txt` file that gives LLMs a small, readable index for a website at inference time. **Does llms.txt help Google rankings?** No. Google says it does not affect Search visibility. **Should I also keep robots.txt updated?** Yes. robots.txt still controls crawler access, so both files matter if you want bots to reach the right pages. ## Sources - [The /llms.txt file](https://llmstxt.org/) — llms-txt - [Latest Google Search Documentation Updates](https://developers.google.com/search/updates) — Google - [Overview of OpenAI Crawlers](https://developers.openai.com/api/docs/bots) — OpenAI - [Perplexity Crawlers](https://docs.perplexity.ai/docs/resources/perplexity-crawlers) — Perplexity - [How Google Interprets the robots.txt Specification](https://developers.google.com/search/docs/crawling-indexing/robots/robots_txt) — Google - [Strand CMS on GitHub](https://github.com/BowTiedSwan/strand) — BowTiedSwan --- # Generative Engine Optimization Explained for AI Search *By The Strand CMS Team · 2026-07-20 · 7 min read* Canonical: https://www.strandcms.com/blog/geo-explained > **Summary:** GEO is the work of making content easy for answer engines to fetch, parse, trust, and cite. It is not a ranking trick or a single tag. GEO, or generative engine optimization, is the ordinary work of making content easy for answer engines to fetch, parse, and quote. This sits in the [best CMS for AI search](/blog/best-cms-for-ai-search-visibility) cluster. If SEO is about finding the page, GEO is about getting the page into the answer. ## What GEO actually means GEO is the work of turning a page into something a machine can use without guessing. That usually means: - clear HTML that contains the real answer - stable URLs - clean headings - visible sources - structured data where it actually fits - a short summary that leads with the point That is not glamorous, but it is what gets cited. ## What generative engine optimization does not mean GEO is not: - a single tag - a schema plugin you install and forget - a promise that Google will boost you - a way to skip actual content quality Google even says that `llms.txt` does not affect Search visibility. That one line removes a lot of nonsense from the conversation. The file can still be useful, just not as a ranking hack. ## The three things answer engines need ### 1. They need to fetch the page OpenAI and Perplexity both publish crawler guidance. Google publishes JavaScript rendering guidance. The common takeaway is that the answer should be in the HTML, not trapped behind a client-only UI. ### 2. They need to parse the structure Headings, short sections, and one topic per page help a model decide what the page is about. That is why answer-first writing matters. It reduces the amount of work needed to identify the core claim. ### 3. They need to trust the claim Visible sources help here. Google's people-first guidance is still the cleanest statement of the rule: write for users, keep the page useful, and do not build around tricks. ## Why Strand leans into GEO Strand does not treat GEO as a layer bolted onto a generic CMS. It bakes the idea into the publishing model. The content lives as MDX in Git, the site can emit a clean Markdown twin, and the plan includes `llms.txt` plus a schema that blocks malformed content before it ships. That matters because GEO is mostly a content operations problem. If the source is clean, the output usually is too. See [Introducing Strand CMS](/blog/introducing-strand) if you want the version that is actually built around that idea. ## A plain-English rule If a smart assistant had to summarize your page in one paragraph, would it get the point right on the first try? If the answer is no, the page probably needs work. ## A simple GEO checklist - Put the answer near the top. - Keep the content in crawlable HTML. - Use headings that mean something. - Add sources where facts appear. - Keep the summary short enough to lift. The goal is not to win a buzzword contest. The goal is to make the site useful enough that an answer engine can quote it without mangling the result. ## A practical workflow A useful GEO workflow is not complicated. Write the page. Add the sources. Check the HTML. Verify that the main answer is visible without the app shell. Then read the page out loud and ask whether the short version still makes sense. If the answer changes once you strip the adjectives away, the page probably needs another pass. ## GEO versus classic SEO Classic SEO still matters because crawlers and search engines still need a discoverable page. GEO adds a second question: once the page is found, can an answer engine extract a trustworthy summary from it? That is why the same page needs crawlable HTML, a clean heading structure, and visible sources. Search visibility gets the page discovered; GEO makes it easier to quote. ## How to prioritize your time If you are deciding where to spend effort, fix the page shape before you reach for schema tricks or helper files. A clear answer near the top, a tight summary, and links to primary sources will do more for citation quality than a vague layer of AI-optimization language. A useful GEO workflow is not complicated. Write the page. Add the sources. Check the HTML. Verify that the main answer is visible without the app shell. Then read the page out loud and ask whether the short version still makes sense. If the answer changes once you strip the adjectives away, the page probably needs another pass. ## Why Strand keeps showing up here Strand is a good example because the publishing model and the machine-readable surface are close together. The article is MDX in Git, the output is meant to stay clean, and the validation rules keep the page from drifting into mush. That is the practical form of GEO. It is not a special trick. It is a habit. ## What a good GEO page looks like A good GEO page usually feels almost plain. The title says exactly what the page is. The first paragraph gives the answer. The next sections add the detail. The FAQ uses the words readers would actually type into a search box. The sources point to the documents that support the claims. That simplicity is the point. It makes the page easier to maintain and easier to cite. ## Why this matters A page that is built for GEO does not have to be loud. It has to be legible. That is a better standard than chasing some imaginary formula because it survives changes in crawler behavior, content tools, and whatever acronym the market comes up with next. If you keep the page clean, a model can pull the answer without rewriting the whole thing. If you keep the sources visible, a reader can check the claim. If you keep the structure stable, the page is easier to maintain after the first publish. That is the work. ## A final editing pass If you are trying to make a real page more GEO-friendly, start with the answer paragraph and the headings. Those are the parts that do the most work for both readers and answer engines. Then check the sources. If the claim is factual and the source is missing, add it. If the page has no real source, cut the claim or rewrite it as opinion. Finally, look at the summary. If the summary reads like a slogan, replace it with a sentence that actually says what the page is about. One more pass usually helps. Read the page out loud and ask whether the first paragraph still works if you remove the adjectives. If the answer gets weaker, the page still needs attention. ## What a finished GEO page looks like If you are trying to make a real page more GEO-friendly, start with the answer paragraph and the headings. Those are the parts that do the most work for both readers and answer engines. Then check the sources. If the claim is factual and the source is missing, add it. If the page has no real source, cut the claim or rewrite it as opinion. Finally, look at the summary. If the summary reads like a slogan, replace it with a sentence that actually says what the page is about. ## FAQ ### Is GEO replacing SEO? No. SEO still matters. GEO is the part that focuses on AI answers and citations. ### Does a better schema block fix GEO? Not by itself. Schema helps, but only when the page already has useful content and a clear structure. ### Is llms.txt enough on its own? No. It can help with discovery, but it does not replace crawlable HTML or good editorial structure. ## Sources - [The /llms.txt file](https://llmstxt.org/) - [Google's Guide to Optimizing for Generative AI Features on Google Search](https://developers.google.com/search/docs/fundamentals/ai-optimization-guide) - [Latest Google Search Documentation Updates](https://developers.google.com/search/updates) - [Creating Helpful, Reliable, People-First Content](https://developers.google.com/search/docs/fundamentals/creating-helpful-content) - [Overview of OpenAI Crawlers](https://developers.openai.com/api/docs/bots) - [Perplexity Crawlers](https://docs.perplexity.ai/docs/resources/perplexity-crawlers) - [Strand CMS on GitHub](https://github.com/BowTiedSwan/strand) ## FAQ **What does GEO stand for?** Generative Engine Optimization. It is the practice of making content easier for AI search systems to read, summarize, and cite. **Is GEO the same as SEO?** No. SEO tries to win the result page. GEO tries to make the page legible enough to be cited inside AI answers. **Does llms.txt improve Google rankings?** Google says it does not affect Search visibility, so no. Use it for machine guidance, not as a ranking lever. ## Sources - [The /llms.txt file](https://llmstxt.org/) — llms-txt - [Google's Guide to Optimizing for Generative AI Features on Google Search](https://developers.google.com/search/docs/fundamentals/ai-optimization-guide) — Google - [Latest Google Search Documentation Updates](https://developers.google.com/search/updates) — Google - [Creating Helpful, Reliable, People-First Content](https://developers.google.com/search/docs/fundamentals/creating-helpful-content) — Google - [Overview of OpenAI Crawlers](https://developers.openai.com/api/docs/bots) — OpenAI - [Perplexity Crawlers](https://docs.perplexity.ai/docs/resources/perplexity-crawlers) — Perplexity - [Strand CMS on GitHub](https://github.com/BowTiedSwan/strand) — BowTiedSwan --- # Get Cited by ChatGPT: Blog Citation Playbook for 2026 *By The Strand CMS Team · 2026-07-20 · 9 min read* Canonical: https://www.strandcms.com/blog/get-cited-by-chatgpt-perplexity > **Summary:** Make the page easy to crawl, easy to render, and easy to summarize. That is what helps ChatGPT, Perplexity, and AI Overviews cite it. If you want to get cited by ChatGPT, Perplexity, or Google AI Overviews, do not start with tricks. Start with the page itself. The page has to be crawlable, the main answer has to be visible early, and the structure has to make it obvious what the article is about. This sits in the [best CMS for AI search](/blog/best-cms-for-ai-search-visibility) cluster. That sounds boring because it is boring. It also works. ## Get cited by ChatGPT and AI Overviews - Put the actual answer in the HTML, not just behind client-side rendering. - Let the bots fetch the page. Check `robots.txt`. - Use one clear primary topic per page. - Add visible sources for claims that can be checked. - Keep the layout simple enough that a machine can summarize it without work. ## 1. Make the page crawlable Google says JavaScript-heavy pages are risky if the content is not present in the initial HTML response. OpenAI publishes crawler guidance, and Perplexity does the same. The common thread is simple: if the main answer is hard to fetch, you are making the crawler do extra work for no reason. For published articles, the safe default is server-rendered or static HTML for the content itself. ## 2. Check your bot access `robots.txt` is not glamorous, but it is where a lot of this starts. Google documents how it interprets the file, and Perplexity says its crawler rules can be managed with robots tags. If you are blocking bots by accident, nothing else in the post matters. A quick check like `curl -s https://example.com/robots.txt` tells you more than a week of guessing. ## 3. Put the answer near the top Answer-first writing helps humans and machines. Start with the direct answer, then explain the edge cases, then give the details. The earlier the answer appears, the less likely a crawler is to skip over it. This is one place where boring writing wins. Clean prose is easier to quote than clever prose. ## 4. Use real sources A claim with a visible source has a better chance of being repeated accurately. Google's helpful-content guidance still points in the same direction: write for people first, not for tricks. That means visible citations, not vague assertions. If you say a crawler supports something, link the crawler docs. If you say Google changed a policy, link the Google note. If you say a result is common, either show the evidence or soften the wording. ## 5. Give the page a simple shape A good AI-citable page usually has the same skeleton: 1. Short answer at the top. 2. A few H2 sections that each cover one idea. 3. A FAQ that matches real search questions. 4. Visible sources. That is also the shape we use in [Strand CMS](/blog/introducing-strand), because it keeps the post easy to maintain after the first draft is gone. ## Why this matters OpenAI publishes bot docs. Perplexity publishes crawler docs. Both are telling you they want to fetch public pages and make sense of them. They are not asking for keyword stuffing. They are asking for content they can parse. That is why the page format matters as much as the topic. A thin page with a good title still loses to a page that actually answers the query. ## A simple checklist - Is the main answer visible without JavaScript? - Is the page allowed in `robots.txt`? - Does the post have one clear topic? - Are the facts cited? - Can a reader copy the summary without confusion? If you can say yes to all five, you are in decent shape. ## A page recipe that gets cited The best pages do not try to sound bigger than they are. They open with the answer, then give the evidence, then give the reader a path to keep going if they want more detail. A useful trick is to write the article as if you were handing it to another editor who only has thirty seconds. Would they know the point from the first paragraph? Would they know what the page proves? Would they know which claims are sourced and which ones are opinion? If the answer is yes, you are close. That same pattern helps on the web because answer engines tend to prefer pages that are easy to slice into small pieces. Short sections are easier to quote than long, circular paragraphs. ## Common mistakes Most weak pages fail for the same reasons. They lead with the brand instead of the answer. They use a catchy heading that does not mean anything. They bury the useful paragraph halfway down the page. They add an FAQ that repeats the title instead of answering actual search questions. The fix is usually to cut. Remove the intro that says nothing. Move the direct answer up. Turn the rest into support material instead of filler. ## A stronger structure A page that gets cited usually has a simple skeleton. The first paragraph answers the question. The next section explains the constraint or the tradeoff. A later section gives the reader something concrete to do. The FAQ closes the loop by answering the questions that keep coming up. That structure is useful because it keeps the page honest. You can see the claim, you can see the proof, and you can see where the article ends. That makes it easier for a reader to trust the page and easier for a model to summarize it without mangling the point. If you want to improve an older post, do not try to save every sentence. Move the parts that help the reader. Cut the parts that only help the writer feel important. ## If you use Strand Strand is built for exactly this workflow. The post lives in Git, the structure is validated, and the final page keeps the article close to its source. That makes it easier to keep the answer stable while the rest of the site changes. [Introducing Strand CMS](/blog/introducing-strand) shows the publishing model in practice. ## Why this matters The page is easier to cite when the answer is obvious, the structure is calm, and the sources sit next to the claims they support. That sounds almost too simple, but the simplicity is what makes the content portable. A model can only quote what it can see, and it can only trust what the page makes clear. If you are writing for a real audience, this usually means dropping the long scene-setting opener and moving the useful paragraph up. It means using headings that describe the actual section, not the mood of the section. It means linking the source instead of writing around the source. None of that is fancy. All of it helps. ## What to do first Take one paragraph at a time and ask whether it earns its place. If it only repeats the title, cut it. If it says something the reader already knows, move it down or remove it. If it adds a fact or a source, keep it. Before you publish, make sure every factual claim in the intro can point to one of the listed sources. If it cannot, the page is still too loose. ## A republish checklist Before you ship an article, confirm the answer appears in the HTML response, the section headings describe real topics, the FAQ answers actual search questions, and the source list contains the documents you used. If any one of those is missing, the page is harder to quote accurately. ## What not to chase Do not chase a fake formula. OpenAI, Perplexity, and Google all point at fetchability, structure, and helpfulness. That means the boring basics matter more than hacks: server-rendered content, clear summaries, and links to primary sources. ## A sensible editorial pattern Write one paragraph that answers the question, one section that explains the constraint, one section that shows the fix, and one FAQ that matches the real follow-up questions. That shape is easier to maintain than a long, clever essay, and it is much easier for an answer engine to lift without mangling the point. ## The last 10 minutes Take a screenshot or open page source, compare the rendered text with the HTML, and make sure the first 100 words still contain the primary keyword. Then check the sources against the exact claim they support. That quick pass catches most citation problems before they ship. If you already have a post and want to improve it, start with the first paragraph and the first two subheads. Those are the spots that usually hide the most friction. Make the answer visible there, then check whether the rest of the page repeats or supports that answer. Then look at the FAQ. If the questions do not match what readers actually ask, rewrite them. If the answers are vague, fix them or remove them. A short FAQ that answers real questions is better than a longer one that exists only for schema. ## A practical editing pass Take one paragraph at a time and ask whether it earns its place. If it only repeats the title, cut it. If it says something the reader already knows, move it down or remove it. If it adds a fact or a source, keep it. That pass usually reveals the real problem. The page is not missing information. It is missing order. ## If you use Strand Strand is built for exactly this workflow. The post lives in Git, the structure is validated, and the final page keeps the article close to its source. That makes it easier to keep the answer stable while the rest of the site changes. [Introducing Strand CMS](/blog/introducing-strand) shows the publishing model in practice. ## Why this matters Most weak pages fail for the same reasons. They lead with the brand instead of the answer. They use a catchy heading that does not mean anything. They bury the useful paragraph halfway down the page. They add an FAQ that repeats the title instead of answering actual search questions. The fix is usually to cut. Remove the intro that says nothing. Move the direct answer up. Turn the rest into support material instead of filler. ## FAQ ### Does a blog need llms.txt to be cited? No. Google says `llms.txt` does not affect Search visibility. It can help as a machine-readable index, but it is not the main event. ### Is SEO still enough? It gets you part of the way there. AI search also cares about how the page is structured, rendered, and sourced. ### Should I hide details to keep the answer short? No. Short answers are fine, but the page still needs enough context for a machine to trust the result. ## Sources - [Overview of OpenAI Crawlers](https://developers.openai.com/api/docs/bots) - [Perplexity Crawlers](https://docs.perplexity.ai/docs/resources/perplexity-crawlers) - [Understand JavaScript SEO Basics](https://developers.google.com/search/docs/crawling-indexing/javascript/javascript-seo-basics) - [How Google Interprets the robots.txt Specification](https://developers.google.com/search/docs/crawling-indexing/robots/robots_txt) - [Creating Helpful, Reliable, People-First Content](https://developers.google.com/search/docs/fundamentals/creating-helpful-content) - [The /llms.txt file](https://llmstxt.org/) - [Strand CMS on GitHub](https://github.com/BowTiedSwan/strand) ## FAQ **What is the fastest way to get cited by AI search?** Put the answer near the top, keep the page crawlable, and make the source clear enough that an engine can quote it without guessing. **Do I need llms.txt for citations?** No. Google says llms.txt does not affect Search visibility. Use it as a helper, not as the main plan. **Do JavaScript-heavy pages lose citations?** Often, yes. If the core answer only appears after client-side rendering, some crawlers will miss it or see it late. ## Sources - [Overview of OpenAI Crawlers](https://developers.openai.com/api/docs/bots) — OpenAI - [Perplexity Crawlers](https://docs.perplexity.ai/docs/resources/perplexity-crawlers) — Perplexity - [Understand JavaScript SEO Basics](https://developers.google.com/search/docs/crawling-indexing/javascript/javascript-seo-basics) — Google - [How Google Interprets the robots.txt Specification](https://developers.google.com/search/docs/crawling-indexing/robots/robots_txt) — Google - [Creating Helpful, Reliable, People-First Content](https://developers.google.com/search/docs/fundamentals/creating-helpful-content) — Google - [The /llms.txt file](https://llmstxt.org/) — llms-txt - [Strand CMS on GitHub](https://github.com/BowTiedSwan/strand) — BowTiedSwan --- # Structured Data AI Search: Speakable, FAQPage, and E-E-A-T *By The Strand CMS Team · 2026-07-20 · 9 min read* Canonical: https://www.strandcms.com/blog/structured-data-for-ai-citations > **Summary:** Schema helps answer engines extract the right parts of a page, but the page still needs clear writing, visible sources, and real editorial quality. Structured data helps machines understand a page, but it does not save a weak page. If the writing is vague, the sources are missing, or the answer is buried, schema will not fix that. This sits in the [best CMS for AI search](/blog/best-cms-for-ai-search-visibility) cluster. If the content is not clear without the markup, the markup has nothing to save. The useful way to think about it is this: schema is a label, not the product. ## Structured data AI search basics The schema types worth paying attention to for AI citations are usually the boring ones: - `FAQPage` for real questions and answers - `SpeakableSpecification` for short passages that should be read aloud or extracted cleanly - article or blog metadata that makes the page easy to identify That is useful, but only when the page itself is worth citing. ## A simple schema stack If you are deciding what to add, start small. Use `FAQPage` when the page really answers common questions. Use `SpeakableSpecification` only when a short passage makes sense as a spoken or extracted summary. Keep the author, article type, and canonical URL honest. That is enough for most pages. A large schema graph is not automatically better. It is just more code to keep accurate. ## Pair markup with editorial work Schema works best when the editorial work is already done. That means the answer appears early. That means the page has visible sources. That means the FAQ questions sound like real search queries. That means the page does not bury the useful part under a brand pitch. If you are building this as a system, the pattern in [Introducing Strand CMS](/blog/introducing-strand) is the simplest version: structure first, markup second. ## When not to use more markup Do not add schema because you feel behind. Do not add FAQ rows just to fill space. Do not add speakable markup to a page that has no short answer section. Do not use structured data as a substitute for sourcing. If the page is already clear, schema can help a machine find the important pieces faster. If the page is not clear, the fix is editorial, not technical. ## A simple working example Imagine a page about a feature launch. The article starts with one sentence that names the feature and the outcome. The next section explains how it works. The FAQ answers the three questions sales keeps hearing. The sources point to the product docs and the policy page. That is a page worth marking up. Now imagine the opposite. The page starts with a slogan, the FAQ repeats the slogan in question form, and the sources section is empty. That page does not need more schema. It needs a rewrite. ## A simple schema workflow Start with the page itself. Write the answer. Add the sources. Decide whether FAQPage or SpeakableSpecification actually fit the content. Then check the markup against the writing instead of treating the markup as a separate project. A useful workflow is to keep the markup small enough that the editor can understand it in one glance. If the schema only makes sense to the person who wrote it, the next edit is going to be painful. ## A deployment checklist Before you ship schema on a real page, read the page without the JSON-LD in your head. Does the answer still make sense? Do the headings still point the reader in the right direction? Are the sources still visible and relevant? If the answer is no, fix the article first. Then check whether the markup adds anything useful. If it helps a crawler understand the page faster, keep it. If it exists only because the template says so, cut it back. That last step matters. Schema is easiest to maintain when it stays small and honest. ## How to choose the right markup Use `FAQPage` when the page answers repeated questions that real readers ask. Use `SpeakableSpecification` only when there is a genuinely short answer block that is worth extracting. Keep the article metadata honest: title, author, canonical URL, and publish date should all match the visible page. When in doubt, default to less. A smaller schema set that matches the prose is more useful than a big graph that drifts out of date. The goal is to help a crawler understand the page, not to build a trophy cabinet of JSON-LD types. ## Why the FAQ change matters Google's removal of FAQ rich results changes the incentive: add `FAQPage` only because it helps structure the page, not because you expect a visual reward in Search. That is a healthier model for AI citation work too, because answer engines are looking for clarity, not decoration. ## What schema can help with Schema can help crawlers identify the page type, the main answer block, and the FAQ. That can reduce ambiguity, especially on pages that also have navigation, cards, or product copy. But it does not replace the editorial work of writing a useful answer, showing sources, and keeping the summary concise. ## How Strand handles it Strand keeps FAQ and sources in the content model so the editor cannot accidentally publish a page that looks structured but reads thin. That makes the schema easier to keep aligned with the prose and gives you a smaller surface to audit before publishing. ## A production review loop Before you ship schema on a real page, compare the visible page and the JSON-LD side by side. The headings should describe the same ideas the schema claims to describe. The FAQ questions should be the exact questions a reader would ask. The sources should support the factual claims, not just decorate the page. If any of those drift apart, fix the article rather than adding more markup. A good production page usually has one schema type doing most of the work, not five. The point is to help a machine find the useful bits faster while keeping the page readable to a human. That makes maintenance easier and keeps the markup honest when the content changes later. Then check whether the markup adds anything useful. If it helps a crawler understand the page faster, keep it. If it exists only because the template says so, cut it back. That last step matters. Schema is easiest to maintain when it stays small and honest. ## A real example Do not add schema to every page just because it is possible. Do not turn FAQPage into a filler machine. Do not add speakable markup to text that is not meant to be read or extracted as a short answer. The more generic the markup gets, the less useful it is. ## FAQPage is still useful Google removed the FAQ rich result feature from Search, but that does not make `FAQPage` invalid. It just means you should stop treating FAQ markup like a trick for extra SERP decoration. Use it when the page genuinely answers common questions. Do not bolt on fake questions to make the schema look busy. ## Speakable is narrow `SpeakableSpecification` is not a magic SEO switch either. It is for passages that make sense to surface as spoken content. If a page has a concise summary or a short answer section, that is the kind of place it belongs. If the page is long and messy, adding more markup will not make it better. ## E-E-A-T is editorial E-E-A-T is not a schema type. It is the quality of the page: who wrote it, why it exists, whether it is useful, and whether the claims are supported. That means the basics still matter more than the tag: - clear author identity - visible sources - a real answer near the top - no filler - no invented certainty ## A practical pattern If you want a page that answer engines can use, write it in this order: 1. Short answer. 2. Supporting detail. 3. FAQ. 4. Sources. That pattern is why [Strand CMS](/blog/introducing-strand) bakes FAQ and sources into the content model. The markup works best when the editorial structure is already doing the heavy lifting. ## What to avoid - Do not add FAQ markup to pages that do not answer questions. - Do not use speakable markup everywhere. - Do not treat schema as a replacement for research. - Do not bury the answer under a marketing intro. ## Why this matters Structured data is most useful when it reinforces a page that already works. A crawler can use the markup to understand the page faster, but only if the page itself makes sense. If the content is vague, more JSON-LD will not rescue it. That is why the editorial checklist still matters. The author has to write clearly. The sources have to be visible. The FAQ has to match real questions. The markup only helps when it sits on top of that work. ## A real example A good product page might use FAQPage for questions that sales hears every week, SpeakableSpecification for a short answer block, and ordinary article metadata for the rest. The same page might also link to the main docs page and a pricing page. That is enough. A bad page would add the same schema types while the headline says almost nothing and the body text is just a brand pitch. That page needs editing, not more structure. ## How Strand handles it Strand keeps the structure and the markup close together so the page does not drift as much. That is useful because it makes it harder to publish a malformed article by accident and easier to keep the FAQ and sources aligned with the draft. If you are building a system for AI search, that is the practical win: the markup and the writing stay in sync. ## A deployment checklist Before you ship schema on a real page, read the page without the JSON-LD in your head. Does the answer still make sense? Do the headings still point the reader in the right direction? Are the sources still visible and relevant? If the answer is no, fix the article first. Then check whether the markup adds anything useful. If it helps a crawler understand the page faster, keep it. If it exists only because the template says so, cut it back. That last step matters. Schema is easiest to maintain when it stays small and honest. ## FAQ ### Do schema and citations solve the same problem? No. Schema helps machines classify the page. Citations help them trust the claim. ### Is FAQPage dead? No. The rich result is gone, but the schema type still exists and is still useful for actual FAQ content. ### Should I add structured data before I write the page? No. Write the page first. Then add the markup that matches it. ## Sources - [FAQPage - Schema.org Type](https://schema.org/FAQPage) - [SpeakableSpecification - Schema.org Type](https://schema.org/SpeakableSpecification) - [Latest Google Search Documentation Updates](https://developers.google.com/search/updates) - [Creating Helpful, Reliable, People-First Content](https://developers.google.com/search/docs/fundamentals/creating-helpful-content) - [SEO Starter Guide](https://developers.google.com/search/docs/fundamentals/seo-starter-guide) - [Google's Guide to Optimizing for Generative AI Features on Google Search](https://developers.google.com/search/docs/fundamentals/ai-optimization-guide) - [Introducing Strand CMS: publishing built for agents](https://strandcms.com/blog/introducing-strand) ## FAQ **Does FAQPage still matter if Google no longer shows FAQ rich results?** Yes. The schema is still valid, even though Google removed the FAQ rich result feature from Search. **Is E-E-A-T a schema type?** No. E-E-A-T is an editorial quality idea, not a JSON-LD property. **Should every page get speakable markup?** No. Use it only where the content really has a short section that should be read aloud or extracted cleanly. ## Sources - [FAQPage - Schema.org Type](https://schema.org/FAQPage) — Schema.org - [SpeakableSpecification - Schema.org Type](https://schema.org/SpeakableSpecification) — Schema.org - [Latest Google Search Documentation Updates](https://developers.google.com/search/updates) — Google - [Creating Helpful, Reliable, People-First Content](https://developers.google.com/search/docs/fundamentals/creating-helpful-content) — Google - [SEO Starter Guide](https://developers.google.com/search/docs/fundamentals/seo-starter-guide) — Google - [Google's Guide to Optimizing for Generative AI Features on Google Search](https://developers.google.com/search/docs/fundamentals/ai-optimization-guide) — Google - [Introducing Strand CMS: publishing built for agents](https://strandcms.com/blog/introducing-strand) — Strand CMS --- # SEO is table stakes. GEO is the moat. *By The Strand CMS Team · 2026-06-14 · 1 min read* Canonical: https://www.strandcms.com/blog/seo-is-table-stakes-geo-is-the-moat > **Summary:** GEO needs content in the server's HTML, clean extractable Markdown, structured data, and cited sources - and MDX-in-Git is unusually well suited to all of it. Classic SEO is now the price of entry. The differentiator is whether ChatGPT, Perplexity, Claude, and Google's AI Overviews can read, trust, and cite you. That is **GEO**, and it has concrete requirements. ## Content has to be in the HTML AI crawlers largely do not run your JavaScript. If your article is assembled client-side, it is invisible to them. Static generation or server rendering - content in the initial HTML response - is non-negotiable. This is why Strand CMS keeps the crawlable surface static. ## Give machines clean source MDX is already the format models ingest cleanly. Strand CMS serves a content-negotiated `.md` of every page and publishes `llms.txt`, so an assistant gets your prose, not your markup. ## Structure and citations build trust `FAQPage` and `speakable` schema, author-entity markup, and visible cited sources all raise the odds that an engine quotes you accurately. Cited claims travel further. ## Why MDX-in-Git fits The same files that render your site are already clean, structured, and diffable. GEO is not a bolt-on for Strand CMS; it falls out of the content model. ## FAQ **What is GEO?** Generative Engine Optimization: making your content easy for AI search engines and assistants to ingest, trust, and cite. **Does serving content from an API hurt SEO?** The indexed surface must be in the server's HTML - static or server-rendered, not assembled client-side. A JSON API for app consumers is a separate channel. ## Sources - [llms.txt proposal](https://llmstxt.org) --- # Why MDX-in-Git beats a database for agent content *By The Strand CMS Team · 2026-06-12 · 1 min read* Canonical: https://www.strandcms.com/blog/why-mdx-in-git > **Summary:** Git gives agent-written content a built-in audit log, review workflow, and rollback - and removes a whole class of database operations from the stack. When an agent is the author, the most useful property of your content store is not query flexibility - it is **traceability**. You want to see exactly what changed, review it, and undo it. Git gives you all three for free. ## Every change is reviewable An agent drafts a post, commits it to a branch, and opens a pull request. The diff is the review surface. A human - or another agent - approves or requests changes. Merging is the publish event. ## Rollback is a revert A bad post is not a database migration; it is a `git revert`. History is never rewritten, so the full record of what was published, and when, is always intact. ## No infrastructure to babysit There is no database to provision, migrate, scale, or back up. The repository is the backup. Deploys are static, so the indexed surface is fast and cheap by construction. ## The tradeoff, honestly Files are not the right answer for highly relational, write-heavy applications. A publication is the opposite: read-mostly, append-mostly, and happiest when its output is static. That is exactly where MDX-in-Git wins. ## FAQ **Doesn't a database scale better than files?** For a publication's read-mostly content, static generation from files is faster and cheaper than a database, and the crawlable surface is fully static. **How do agents edit content safely?** Agents write MDX on a branch and open a pull request. CI validates the frontmatter; nothing merges to main directly. ## Sources - [Strand CMS design notes](https://github.com/BowTiedSwan/strand) --- # Introducing Strand CMS: publishing built for agents *By The Strand CMS Team · 2026-06-10 · 1 min read* Canonical: https://www.strandcms.com/blog/introducing-strand > **Summary:** Strand CMS keeps the ~15% of a CMS that earns traffic - a validated content schema and an SEO/GEO core - and replaces the editor with agents and the database with Git. Most content systems were built for a person sitting at an editor. Strand CMS starts from a different premise: the writing, the SEO, the structured data, the scheduling, and the publishing can all be done by an agent working against a well-defined contract. ## What we kept, and what we threw out A modern CMS is mostly machinery for human editors - a WYSIWYG editor, roles and permissions, members and subscriptions, newsletters, a theme marketplace. The genuinely valuable part is small: a validated content schema and a publishing-quality core that produces sitemaps, JSON-LD, canonical tags, and feeds. Strand CMS keeps that core and throws out the rest. The editor becomes a set of agent skills. The database becomes Git. Every article is a commit; publishing is a pull request. ## The schema is the contract Because an agent writes the post, the schema matters more, not less. Frontmatter is validated on commit, so a malformed article literally cannot merge. That validation is the same check an agent runs to correct itself before opening a PR. ## SEO is the floor; AI search is the reason Strand emits sitemaps, JSON-LD, RSS, and robots like any good CMS - and then goes further, with `llms.txt` and a clean Markdown rendering of every page so AI search engines ingest your content instead of your hydrated DOM. ## Try it One command scaffolds a running publication, wired for an agent to take over. The whole thing is open source. ## FAQ **What is Strand CMS?** An open-source publishing system for programmatic blogs and news sites. Articles are MDX files in Git, written by AI agents, with SEO and AI-search built into the core. **How is it different from Ghost or WordPress?** No database, no WYSIWYG editor, no plugin marketplace. The editor is replaced by agents writing against a strict schema; the database is replaced by Git. ## Sources - [Strand CMS on GitHub](https://github.com/BowTiedSwan/strand)