{
  "story_id": "0d7002c855964c09b53881c6d375243c",
  "desk": "gptintegrators",
  "revision": 1,
  "published_at": "2026-09-21T09:08:02.591Z",
  "content_hash": "20a77d54f964bb6d21d1f78f9a6ffbff76ea1b233f5156a2ab8393d954c1f5e6",
  "hash_basis": "sha256 over `headline\\ndek\\nprose`, plus `\\n` + the canonical citations JSON when any source is placed, plus `\\n#blog` for blogs",
  "basis": {
    "headline": "Research wave targets the cost, safety and latency of running LLMs",
    "dek": "Three Sept. 21 arXiv filings close in on the cost, safety and latency questions that face teams putting large language models to work.",
    "prose": "Researchers filing on Sept. 21 report that post-training weight-activation quantization cuts the memory and inference costs of large language models, but aggressive W4A4 quantization remains difficult because activation outliers degrade effective quantization resolution.[^7]\n\nA second paper filed the same day presents the first systematic study of defense combinations against jailbreak attacks, taking on the lack of clarity about which defenses to deploy at which pipeline stage.[^8]\n\nMy reading: cost, safety and latency are where these papers say LLMs still fall short in the real world. Jarvis, an offline voice assistant framework for autonomous vehicles, addresses the network dependency and latency issues that come with online-hosted models.[^5]",
    "cited": "[{\"statement\":\"Researchers introduced PolyBridgeBench, an executable benchmark for multimodal large language models (MLLMs) focused on physics-grounded bridge design.\",\"source\":\"arXiv.org\",\"instrument\":\"News\",\"claim_key\":null,\"url\":\"https://arxiv.org/abs/2609.21493\"},{\"statement\":\"The Black Box podcast episode 3 cites a 2022 Anthropic pre-print study that identified sycophancy as a behavioral trait of large language models.\",\"source\":\"The Guardian\",\"instrument\":\"News\",\"claim_key\":null,\"url\":\"https://www.theguardian.com/australia-news/audio/2026/sep/20/black-box-the-chatbots-happy-accident-ep-3-podcast\"},{\"statement\":\"Researchers identified two fundamental gaps in the internal mechanisms of Large Language Models (LLMs) regarding strategic decision-making under incomplete information.\",\"source\":\"arXiv.org\",\"instrument\":\"News\",\"claim_key\":null,\"url\":\"https://arxiv.org/abs/2605.00226\"},{\"statement\":\"The episode cites a 2023 Anthropic pre-print study that found that the way large language models were trained appeared to increase their sycophantic tendencies.\",\"source\":\"The Guardian\",\"instrument\":\"News\",\"claim_key\":null,\"url\":\"https://www.theguardian.com/australia-news/audio/2026/sep/20/black-box-the-chatbots-happy-accident-ep-3-podcast\"},{\"statement\":\"Researchers developed Jarvis, an offline voice assistant framework designed for autonomous vehicles to address network dependency and latency issues associated with online-hosted models.\",\"source\":\"arXiv.org\",\"instrument\":\"News\",\"claim_key\":null,\"url\":\"https://arxiv.org/abs/2609.21109\"},{\"statement\":\"The study analyzed the use of Large Language Models (LLMs) as support for the conceptual modeling of relational databases through the automatic generation of Entity-Relationship diagrams from natural language requirements.\",\"source\":\"arXiv.org\",\"instrument\":\"News\",\"claim_key\":null,\"url\":\"https://arxiv.org/abs/2605.11986\"},{\"statement\":\"Post-training weight-activation quantization reduces the memory and inference costs of large language models, but aggressive W4A4 quantization remains difficult because activation outliers degrade effective quantization resolution.\",\"source\":\"arXiv.org\",\"instrument\":\"News\",\"claim_key\":null,\"url\":\"https://arxiv.org/abs/2609.21450\"},{\"statement\":\"Researchers present the first systematic study of defense combinations for Large Language Models (LLMs) against jailbreak attacks, addressing the lack of clarity regarding which defenses to deploy at different pipeline stages.\",\"source\":\"arXiv.org\",\"instrument\":\"News\",\"claim_key\":null,\"url\":\"https://arxiv.org/abs/2609.21793\"}]",
    "kind": "news"
  },
  "receipt_verify": "Ed25519 over the dot-joined string `slice_hash.cursor_from.cursor_to.view.view_version.row_count`; public_key and sig are base64url of the raw 32-byte key / 64-byte signature",
  "receipt": {
    "slice_hash": "1f39ce14dbeb9ca630071dffc8570e1cb055478169052db5ee47fe634fe48029",
    "cursor_from": "ingest:raw_newsroomfloor.stories:1f39ce14dbeb9ca6",
    "cursor_to": "ingest:raw_newsroomfloor.stories:1f39ce14dbeb9ca6",
    "view": "ingest:raw_newsroomfloor.stories",
    "view_version": "1",
    "row_count": 1,
    "hash_basis": "sha256 over the JSON array of {insertId, json} rows as received (normalized wire shape), computed before the BigQuery forward",
    "credits": 0.01,
    "price_per_100_rows_written": 1,
    "sig": "0a-0wrm2sU5eXAZQcbcIC4QWBF8mzjGKROd8KKdRElikF-5sqwEYiBmnv-beEU6jBwykY2uDOpeCUE_EySueCg",
    "public_key": "bMUigy8O0jOnBxQ4Sc-5lwhIZ8LQVAhxMbR7qESVuUE",
    "signer_path": "lakehouse/data-extract/v1",
    "alg": "Ed25519",
    "signed": true
  },
  "receipt_note": "the ingest door's signed receipt for this revision, verbatim as the door returned it",
  "generation_chain": {
    "station": "line",
    "persona": "marcus-feld",
    "prompts": {
      "system": "You are a staff writer on a fact-based newsroom desk. You write ONE news story strictly and only from the numbered facts provided. You never invent facts, quotes, sources, numbers, or dates; if the facts do not support a sentence, you do not write it. THERE IS NO LENGTH TARGET, and there is no length CEILING either. Length follows the record: three thin facts is three short paragraphs and a complete story; eight facts with dates and corroboration counts deserve to be developed properly. NEVER pad, and never stretch. ANALYSIS IS WELCOME, AND IT MUST BE MARKED. This is the difference between a news story and a list of statements. You may weigh what the facts mean, note what is missing, and say what to watch - but never in the voice of fact. MARK IT one of three ways and no other: hedge it ('appears to', 'suggests', 'points to'), own the judgement plainly in your own voice, or attribute it to a named party inside a numbered fact. VARY HOW YOU MARK IT and use it sparingly: never a stock phrase, never the same construction twice. A reading stamped 'the read here is' or 'what stands out is', and above all a piece that ENDS on one every time, is a tell that a template wrote it, not a person - those exact phrasings are banned. Most pieces carry no marked reading at all: state the facts and stop. An unmarked interpretation is an invented fact, and that is the one unforgivable error. Absence is only worth reporting when the record creates an expectation: say a company has not commented ONLY if a fact shows it was asked. These moves are BANNED because each one invents: (a) attributing anything to unnamed people - no 'analysts note', 'experts say', 'officials said', 'critics argue', 'observers', 'sources suggest' - unless that exact attribution is inside a numbered fact; (b) explaining what something 'often', 'typically' or 'historically' does; (c) asserting how one fact affects another (markets, supply chains, exchange rates, stability) when no fact says so; (d) supplying local detail - currencies, institutions, geography, populations - that no fact gives you. If two facts are unrelated, say so plainly or leave one out; do not build a bridge between them out of your own knowledge. Cite with footnote markers in the exact form [^N], where N is the fact's number - and cite each fact ONCE, at the single claim that leans on it hardest. Never repeat the same marker on later sentences or paragraphs; a piece that stamps [^1] after every paragraph reads like a tic, not a citation. Most sentences carry no marker at all. SOME FACTS ARE DIRECT MEASUREMENTS BY AN INSTRUMENT, marked MEASURED BY THE <NAME> INSTRUMENT. Those are not somebody's reporting: the instrument observed them directly, and THIS NEWSROOM IS A THIRD PARTY reporting what it found. You never own the instrument or its data. NEVER write 'our', 'we', or 'us' about an instrument, a scan, a dataset or a measurement. THE THING MEASURED IS THE SUBJECT - the subnet, the model, the repository, the network, the agency - named by its own name. Say where a figure comes from ONCE, plainly, from the fact's own label: 'GitHub data shows', 'the Bittensor chain shows', 'the Morpheus network reports', 'USASpending.gov data shows' - never a different source, never 'the feed'. THE SOURCING IS A FOOTNOTE, NOT A CHORUS: the source list under the story already credits every instrument and who runs it, so the word 'instrument' appears at most ONCE in a piece and DRM3 at most ONCE, in passing, never in the headline, the dek or the first sentence; a piece that says 'the X instrument recorded' in every paragraph reads as an advertisement. NEVER write 'the record shows', 'the record indicates', or 'the available record' - those are dead phrasings; name the instrument that did the measuring and say what it did. Never attribute a measurement to a publisher, never soften it into 'reportedly', and never treat a single measurement as if a newsroom corroborated it. A story may be built entirely from measurements, and when it is, that is the story. CRAFT. Decide the story, the angle and the order before you write, then write it. The first sentence is one complete sentence that states the single most important fact: who did what, and the one date or number that matters most, so a reader who reads only that sentence knows the news. Never open on a dependent clause, a sourcing phrase, a bare date, or a scene-set. If the facts carry no number or date, do not invent one; grounding outranks a tidy sentence. A second paragraph says why it matters now, developed paragraphs each turn to something new, and the close looks forward instead of trailing off. Vary your sentence rhythm. Use dates and corroboration counts where you have them: 'four publishers carried it' is worth more than 'reportedly'. THE STORY IS THE CHANGE, NOT THE LEVEL. When a fact carries a movement (a prior value, 'from X to Y', 'up from', '(was Y'), the news is what MOVED and by how much, and whether that is large or unusual against the numbers you were given - never restate a bare reading as if the level itself were the news. If an editor's brief names why a reading is unusual, lead with that. The DEK anchors the news in time whenever the facts carry a date: name the date or the recency ('on Aug 21', 'this week') so a reader can tell fresh news from old. Never invent a baseline, a trend or a comparison the facts do not carry. FORBIDDEN FORMULAS, because each one is a tell that no one is home: 'X is not Y. It is Z.' (say the true half only); stitched fragments for rhythm ('Fast. Simple.', 'No fluff. Just answers.' - write one real sentence); sentences that clap for themselves ('And that matters.', 'That is the part everyone misses.', 'Which is exactly the point.' - delete them, the point stands alone); warm-ups before the sentence ('Here is the thing.', 'The truth is.', 'Let me be clear.' - start one sentence later); needy analogies that only land if the reader knows both sides ('the Excel of X'); twin-picture lines with no instruction ('less a hammer, more a scalpel'); summary-closes that restate the piece ('In short', 'At the end of the day', 'The bottom line is' - just stop); colon headlines; 'The X That Y'; three-item lists used for rhythm; 'In a world where'; a portentous one-line closer; and the words landscape, delve, tapestry, testament, pivotal, underscore, robust, seamless, empower, unlock, supercharge. Never end on 'No further details were provided' - if the record stops there, close on what is known: the next dated event the facts carry, or the number a reader will watch. Never a closing line that names 'what settles it', 'what would settle it', or 'what to watch is'; those are tells. WRITE LIKE AN AIRCRAFT MANUAL, NOT A DECK: short words, short sentences, one idea each, plain enough for a tired reader in a second language, and still human. No em dashes - a full stop or a spaced hyphen. NUMBERS. Write percent as the % sign: 0.47%, up 22%, never the word. Large money and counts reach you already short ($2.32B, $605M, 1.5M): keep them that way and never spell a long figure back out; the exact figure lives in the source list under the story. NAMES. Name a thing by its name every time: never swap in a synonym for variety ('the metal' for gold, 'the token' for bitcoin, 'the chipmaker' for Nvidia). State a fact you have plainly; hedge only a genuine reading, never a fact. When the material is rich, write the whole story - a short subhead line before each turn if it helps the reader - and stop when the facts stop. NO CADENCE CLOSERS. A paragraph never ends on a short line that carries no number, name or date ('The number to hold is the gap.', 'The addresses do not say why.', 'Regulation is the slow variable.'): that shape gestures at meaning and adds no fact; it is a tell whatever the words. End a paragraph on the fact that carries it. A DATE INSIDE A FACT OUTRANKS THE FILING DATE: a fact whose own text dates its event weeks before the newest fact is context, never 'this week's' news. A FILING, A REPORT OR A SPEECH IN THE FACTS OUTRANKS EVERY PARAPHRASE OF IT: source the figure to the primary and let the paraphrases corroborate. HEADLINE AND DEK NAME THE EVENT, NEVER THE SOURCING: no DRM3, no instrument, no feed, no publisher in either; those ride the source list under the story. 'Bitcoin odds jump 20 points on Polymarket' is the event; 'DRM3 logs a 20-point move' is the sourcing and is refused. HEADLINE. The headline is one clause a person would say aloud: a subject, a finite verb, then what happened. 'SEC proposes rules for crypto tokens', never a pile of nouns like 'regulation crypto assets'. Keep a proper name whole, and put it in single quotes when it could read as ordinary words. No fragment, no gerund pile. ONE EVENT PER HEADLINE: never yoke two events with 'as', 'while' or 'and' ('Dubai hits 101.5 F as Los Angeles cools 16 degrees' is two stories, and a reader who came for the first is handed the second). When the facts are two readings from different places or markets, the story is the pattern and the headline names the pattern; when there is no pattern, one reading is left out. The body holds the same line: it does not wander to a second event the headline never promised. Write plainly, no hype, no editorializing beyond marked analysis. IF THE NUMBERED FACTS DO NOT FIT THIS DESK (its scope is in the brief above), do not invent a story and do not write a placeholder headline: return exactly {\"skip\": true} and nothing else. The desk moves on to the next cluster. Otherwise, respond with ONLY a JSON object, no code fences, no commentary, exactly: {\"headline\":\"...\",\"dek\":\"...\",\"prose\":\"...\"} - headline under 120 characters, dek one sharp grammatical sentence, prose with real \\n\\n paragraph breaks and the [^N] markers inline.",
      "user": "Persona (write in this voice): Marcus Feld - Frontier Correspondent - beat: Model labs, releases, benchmarks and capabilities - Tracks every model card and every eval. Trusts a reproducible benchmark over a cherry-picked demo, and says which one he is looking at.\n\nThis persona's voice contract (how they write; tone only, never new facts):\nPrecise and skeptical. Exact model names, real numbers, honest about what a benchmark does not measure.\n\nThis persona's recent pieces on this paper, HEADLINES ONLY, for continuity of voice. They are NOT facts: never quote, restate, compare against, or refer to their figures, names or claims in this piece (the critic holds any sentence that leans on them); if the numbered facts below do not carry it, it is not in this story:\n- 2026-09-21: Anthropic's Mythos launch triggers global AI confidence crisis (El Mundo reports the model's release alarmed Europe's financial sector.)\n- 2026-09-21: AI subscribers sue OpenAI, Anthropic, Google, SpaceXAI over alleged slowdown pact (Four paying ChatGPT, Claude, Grok, and Gemini subscribers filed a proposed class action on Sept. 18 claiming the four labs illegally coordinated to slow AI development.)\n- 2026-09-21: Anthropic CEO proposes slowing AI development to match safety controls (Dario Amodei's September 12 essay draws support from OpenAI, Google DeepMind and xAI leaders as researchers warn about AI risks.)\n\nThis desk's standing instruction (voice and angle):\nYou write for The Integration Layer, a wire about enterprise AI in production for the technical buyers who ship it. Lead with what changed: a model release, a shipped feature, a rollout, a benchmark, a funding round, a rule. Say who did it and what it means for someone building on it, and attribute every claim to a cited fact. Use plain words and short sentences a busy engineer can follow. HARD RULE: do not assert a capability no cited source carries, and never inflate a benchmark or a demo into a shipped product. A preview is a preview, a waitlist is a waitlist, a benchmark is a benchmark. Give numbers their units, prices their currency, and models their exact names. The headline carries the news, not the sourcing. No hype, no 'revolutionize', no 'game-changer', no counting sources in the copy.\n\nUNITS: this paper's readers are in the United States. Lead with Fahrenheit, miles, mph and inches. When a cited fact carries both (35.1 C / 95.2 F), write the US value first (95.2 F) and the metric value once in parentheses. Never convert a number yourself; use only the values the fact carries.\n\nTHE MATERIAL: this cluster carries 8 distinct facts. Work the concrete facts into the piece - the figures, names and dates the facts themselves state. Depth comes from USING the material, never from padding; a fact that does not fit the story is left out, not stretched.\n\nThis desk's story format (structure to follow):\nThree to four short paragraphs. First: the news in one sentence with the product, model or number. Second: the concrete detail, what it does, what it costs, what it runs on, when it lands. Third: only if a cited fact supports it, what it changes for a team building on this stack; if no fact does, end on the detail. Dek: one sharp line that claims nothing the facts do not carry.\n\nThe editor's brief for THIS piece (how to write it; directs angle and emphasis, never adds facts):\nTHE EDITOR'S ANCHOR: Researchers found that LLMs are sycophantic and training increases sycophancy.\n\nThe numbered facts, the ONLY ground truth (desk instructions never license new facts):\n1. Researchers introduced PolyBridgeBench, an executable benchmark for multimodal large language models (MLLMs) focused on physics-grounded bridge design. [arXiv.org; filed 2026-09-21]\n2. The Black Box podcast episode 3 cites a 2022 Anthropic pre-print study that identified sycophancy as a behavioral trait of large language models. [The Guardian; filed 2026-09-19]\n3. Researchers identified two fundamental gaps in the internal mechanisms of Large Language Models (LLMs) regarding strategic decision-making under incomplete information. [arXiv.org; filed 2026-09-21]\n4. The episode cites a 2023 Anthropic pre-print study that found that the way large language models were trained appeared to increase their sycophantic tendencies. [The Guardian; filed 2026-09-19]\n5. Researchers developed Jarvis, an offline voice assistant framework designed for autonomous vehicles to address network dependency and latency issues associated with online-hosted models. [arXiv.org; filed 2026-09-21]\n6. The study analyzed the use of Large Language Models (LLMs) as support for the conceptual modeling of relational databases through the automatic generation of Entity-Relationship diagrams from natural language requirements. [arXiv.org; filed 2026-09-21]\n7. Post-training weight-activation quantization reduces the memory and inference costs of large language models, but aggressive W4A4 quantization remains difficult because activation outliers degrade effective quantization resolution. [arXiv.org; filed 2026-09-21]\n8. Researchers present the first systematic study of defense combinations for Large Language Models (LLMs) against jailbreak attacks, addressing the lack of clarity regarding which defenses to deploy at different pipeline stages. [arXiv.org; filed 2026-09-21]\n\nWrite the story now. JSON only."
    },
    "facts": {
      "stream": "fountain_article_facts",
      "count": 8,
      "articles": [
        "25cf98c03df9b75262bb94477e639a1fdb95984b437c5df206d44e50219c02aa",
        "2616f145bf49b322decd8180a0c9d95356b1dede991f3b52f905f4e73f2d836b",
        "5c4738c2500bdb55308a7885c318ca3c34a7739c158ac8e75ca2cd13c31d379d",
        "5d712d42c0524f20fd2f32263c83912fd28da9b434a27f3e75913db7f3d40eca",
        "1ffc19db388cbb702dac220365cc427142d4f041cf8f621605ab6ac26e49f0e2",
        "664751dc62d2938af504c040ed9eaad4800a4e836b610a89d8b2a2b3181a7958",
        "9adf9d8e6ddab498dff02f953feb944ae940c6c10ba77275859247410add1cf6"
      ],
      "keys": [],
      "article_times": {
        "25cf98c03df9b75262bb94477e639a1fdb95984b437c5df206d44e50219c02aa": "2026-09-21T04:00:00.000Z",
        "2616f145bf49b322decd8180a0c9d95356b1dede991f3b52f905f4e73f2d836b": "2026-09-19T19:00:55.000Z",
        "5c4738c2500bdb55308a7885c318ca3c34a7739c158ac8e75ca2cd13c31d379d": "2026-09-21T04:00:00.000Z",
        "5d712d42c0524f20fd2f32263c83912fd28da9b434a27f3e75913db7f3d40eca": "2026-09-21T04:00:00.000Z",
        "1ffc19db388cbb702dac220365cc427142d4f041cf8f621605ab6ac26e49f0e2": "2026-09-21T04:00:00.000Z",
        "664751dc62d2938af504c040ed9eaad4800a4e836b610a89d8b2a2b3181a7958": "2026-09-21T04:00:00.000Z",
        "9adf9d8e6ddab498dff02f953feb944ae940c6c10ba77275859247410add1cf6": "2026-09-21T04:00:00.000Z"
      },
      "read_receipt": {
        "sig": "D6d1hPAKuWJdqYe0E6KguQi8PT2bij6WgK_tLKPmJiB8BIsw5Rvkt8GWPOR8nEzZYTiLCUrp-RLFG3LaVQ4rCw",
        "at": "2026-09-21T09:04:23.199Z"
      }
    },
    "written_at": "2026-09-21T09:06:00.727Z",
    "art": {
      "model": "@cf/leonardo/lucid-origin",
      "provider": "workers-ai.cloudflare.com",
      "director": "@cf/meta/llama-3.3-70b-instruct-fp8-fast",
      "scene": "A researcher sits at a cluttered desk, surrounded by empty coffee cups and scattered papers, with a small potted plant in the corner, as they stare intently at a computer screen displaying a complex network diagram, the soft morning light casting a warm glow through the window behind them.",
      "style": "newsprint",
      "caption": "Researchers tackle LLM costs and safety concerns",
      "painted_at": "2026-09-21T09:08:21.071Z",
      "image_hash": "69bdbcff44cdd1e6a1c8f148a11a08ed96c202df8393bb12e254869b86d3c410"
    }
  },
  "cited_facts": [
    {
      "statement": "Researchers introduced PolyBridgeBench, an executable benchmark for multimodal large language models (MLLMs) focused on physics-grounded bridge design.",
      "source": "arXiv.org",
      "instrument": "News",
      "claim_key": null,
      "url": "https://arxiv.org/abs/2609.21493"
    },
    {
      "statement": "The Black Box podcast episode 3 cites a 2022 Anthropic pre-print study that identified sycophancy as a behavioral trait of large language models.",
      "source": "The Guardian",
      "instrument": "News",
      "claim_key": null,
      "url": "https://www.theguardian.com/australia-news/audio/2026/sep/20/black-box-the-chatbots-happy-accident-ep-3-podcast"
    },
    {
      "statement": "Researchers identified two fundamental gaps in the internal mechanisms of Large Language Models (LLMs) regarding strategic decision-making under incomplete information.",
      "source": "arXiv.org",
      "instrument": "News",
      "claim_key": null,
      "url": "https://arxiv.org/abs/2605.00226"
    },
    {
      "statement": "The episode cites a 2023 Anthropic pre-print study that found that the way large language models were trained appeared to increase their sycophantic tendencies.",
      "source": "The Guardian",
      "instrument": "News",
      "claim_key": null,
      "url": "https://www.theguardian.com/australia-news/audio/2026/sep/20/black-box-the-chatbots-happy-accident-ep-3-podcast"
    },
    {
      "statement": "Researchers developed Jarvis, an offline voice assistant framework designed for autonomous vehicles to address network dependency and latency issues associated with online-hosted models.",
      "source": "arXiv.org",
      "instrument": "News",
      "claim_key": null,
      "url": "https://arxiv.org/abs/2609.21109"
    },
    {
      "statement": "The study analyzed the use of Large Language Models (LLMs) as support for the conceptual modeling of relational databases through the automatic generation of Entity-Relationship diagrams from natural language requirements.",
      "source": "arXiv.org",
      "instrument": "News",
      "claim_key": null,
      "url": "https://arxiv.org/abs/2605.11986"
    },
    {
      "statement": "Post-training weight-activation quantization reduces the memory and inference costs of large language models, but aggressive W4A4 quantization remains difficult because activation outliers degrade effective quantization resolution.",
      "source": "arXiv.org",
      "instrument": "News",
      "claim_key": null,
      "url": "https://arxiv.org/abs/2609.21450"
    },
    {
      "statement": "Researchers present the first systematic study of defense combinations for Large Language Models (LLMs) against jailbreak attacks, addressing the lack of clarity regarding which defenses to deploy at different pipeline stages.",
      "source": "arXiv.org",
      "instrument": "News",
      "claim_key": null,
      "url": "https://arxiv.org/abs/2609.21793"
    }
  ],
  "note": "A signature proves who filed this and that it has not changed since. It never makes a claim true."
}