// Knowledge Ingestion Policy · v1.1 · In Force 2026-04-30

How S.A.G.E.
learns from the web.

Version
v1.1
In Force
April 30, 2026
Bot User-Agent
OER-Oracle-Bot/1.0
Contact
opt-out@oneearthrising.com

This page is the public summary of our internal Knowledge Ingestion Policy v1.1. It describes what content we ingest into S.A.G.E. — the knowledge layer that powers our gaming AI companion — under what license terms, and how creators or rights-holders can request removal.

Want your content removed? Submit a request and we'll action it within 7 business days.

▸ OPT OUT

Governing principles

Six principles govern every ingestion decision we make. They are stated here in plain language; the binding internal versions appear in our policy document.

  1. License before access. We do not ingest content unless we have a documented legal basis. "Publicly accessible" is not a license.
  2. Transformation, not reproduction. We embed paraphrased summaries, not raw source text. Verbatim excerpts are stored only as short evidence snippets for citation back to you.
  3. Attribution and link-back are mandatory. Every chunk of knowledge in S.A.G.E. carries a source URL. S.A.G.E. cites its sources visibly to users.
  4. Creator priority. When given a choice between ingesting a creator's content without permission and inviting them into our paid Knowledge Keepers program, we choose the program. Always.
  5. robots.txt and rate limits are non-negotiable. Any source that prohibits crawling is excluded. We never burden a host.
  6. Opt-out is honored. Anyone who requests removal has their content purged within 7 business days, with confirmation.
▸ Architectural Firewall · New in v1.1

Retrieval is not training.

Ingested content is used exclusively for retrieval: chunks are fetched at the moment a user asks a question, surfaced with citation, and contribute to S.A.G.E.'s response only as cited evidence.

Ingested content is never used to fine-tune, distill, or otherwise modify the parameters of the underlying language model that S.A.G.E. depends on. The embedding pipeline writes to S.A.G.E.; no pipeline writes to the model. This firewall is architectural, not aspirational.

The legal basis for retrieval (transformative quotation with attribution) is materially different from the legal basis for training, and we keep them separate by design.

What we ingest, and what we don't

Sources are sorted into categories by their legal basis. Only Categories A, B, and C may be ingested. Category D may not.

Permitted

Wikipedia and other Wikimedia properties (CC-BY-SA), public-domain government and academic sources, content explicitly published under Creative Commons CC-BY or CC0. Accessed through official APIs. Attribution preserved on every chunk.

Permitted

Sources accessed through official APIs whose terms permit our use: YouTube Data API (metadata), Reddit Data API (commercial license), Steam Web API, IGDB, Twitch API, GiantBomb. Plus game publishers' own documentation, patch notes, and developer blogs where their public terms permit quotation with attribution.

Permitted (preferred)

Content submitted by creators through our Knowledge Keepers program. The contributor retains copyright; we receive a non-exclusive license to use the content within S.A.G.E. and pay the contributor per query citation. This is our preferred path whenever Categories A and B don't cover what we need.

Not permitted

Out of scope: YouTube transcripts obtained outside the Data API; Reddit content scraped outside the licensed Data API; Fandom/Wikia pages where terms prohibit AI ingestion; sites whose robots.txt or terms of service prohibit scraping or AI use; gaming forum user-generated content (GameFAQs, Steam community guides, Discord) without explicit author permission; paywalled content; anything containing third-party personal data.

Category D sources never enter S.A.G.E. Our automated pipeline is configured to fail closed — if classification or robots.txt enforcement fails, the run aborts.

Signals we respect

Our crawler identifies itself as OER-Oracle-Bot/1.0 (+https://oneearthrising.com/sage-policy) on every request. Before fetching any URL it reads the following signals. If any tell us to stay out, we stay out.

Rate limits are also non-negotiable: no more than 1 request per 2 seconds per host, no parallel requests to the same host. We monitor emerging standards quarterly and commit to adopting new signals within 60 days of stabilization.

How we process content

What lands in S.A.G.E. is not a copy of what we fetched. Every approved source goes through structured extraction before any chunk is stored:

4.1 Time-to-live for Category B sources

Category B sources accessed through licensed APIs are subject to a 90-day time-to-live. Before TTL expires, the source is re-verified against the live API. If the source content has been deleted, materially altered, or made private since ingestion, the corresponding chunks are purged from S.A.G.E. immediately. This honors pass-through deletion requirements from API providers.

Category A and Category C sources are exempt from strict TTL deletion unless an explicit opt-out is received.

4.2 Factual correction

Even legally acquired content can become factually incorrect — through wiki vandalism prior to revert, or through stale information after a game patch. Operators and trusted graders can flag chunks for re-extraction or quarantine. Flagged chunks are excluded from retrieval pending review.

How to opt out

You can request removal of your content from S.A.G.E. at any time. There are three paths, depending on what fits you best:

01.
Self-service form — fill out our opt-out request form. You'll get a tracked reference number.
02.
Email — write to opt-out@oneearthrising.com with the URL or domain you'd like us to stop indexing.
03.
Machine-readable signals — add any of the signals listed in §3 to the URL or domain you control. We'll honor it on the next crawl.

Our service-level commitment for opt-out requests:

2
Business days to acknowledge
7
Business days to purge content
Written confirmation on removal

Opt-out applies to future ingestion as well: once a domain or creator handle is on our opt-out registry, it is excluded from all subsequent runs automatically.

How S.A.G.E. uses it

S.A.G.E. is configured to synthesize, abstract, and cite — never to reproduce. Specifically:

Legacy content — reclassification complete

All content originally tagged legacy_pre_policy has completed CEO-led reclassification review. Every chunk currently in S.A.G.E. was confirmed to trace directly to a named Knowledge Keeper submission and has been reclassified into Category C (creator-licensed) — our preferred path per §2. No content required quarantine. The legacy_pre_policy tag no longer applies to anything in the corpus.

Live, current figures are always available on the Knowledge Source Transparency Report →

Transparency & EU compliance

The legal landscape for AI training data is unsettled and moves quickly. We adopt the more conservative interpretation when an obligation is ambiguous — the cost of over-disclosure is small, the cost of under-disclosure is large.

Knowledge Source Transparency Report

We publish a live, self-updating Knowledge Source Transparency Report aggregating, per game: license categories and chunk counts. The report covers all content ingested under this policy (legacy content joins once reclassification completes). View the live report →

We no longer crawl third-party platforms directly. All new knowledge enters S.A.G.E. through direct collaboration with Knowledge Keepers — creators who submit their own content and are paid per citation. Learn more or apply on the Knowledge Keepers overview page.

EU AI Act

The EU AI Act (Article 53) imposes transparency obligations on providers of General-Purpose AI (GPAI) models. Whether S.A.G.E. constitutes a GPAI model under the Act is not yet settled — a defensible position is that S.A.G.E. is a retrieval-augmented application built on top of Google's Gemini (the GPAI provider), not a GPAI model itself. We are pursuing qualified EU legal review and will revise this policy based on counsel's recommendations.

▸ Note

This policy is not legal advice. It reflects best practice as understood by the CEO and an AI strategic advisor as of April 30, 2026. We will obtain qualified legal review of this policy and the Knowledge Keepers license terms before any production rollout beyond closed beta.

Updates to this policy

This is version 1.1, in force April 30, 2026. New in this version, marked above: machine-readable rights reservations beyond robots.txt; 90-day TTL for Category B; factual correction surface; the architectural firewall between retrieval and training; EU AI Act considerations; and signature/changelog discipline.

All v1.1 changes are forward-compatible with v1.0 — no v1.0 commitment is removed. Material changes are versioned, dated, and summarized at the top of this page. Previous versions remain accessible by request.

Questions? Email opt-out@oneearthrising.com for opt-out matters, or hello@oneearthrising.com for general inquiries.