Analysis of 22 GB Codex conversation data finds 71% of storage is screenshots

A developer analyzed their complete Codex conversation history totaling 22.83 GB across 1,231 files and found that embedded images account for 71% of the storage footprint. The byte-level analysis, conducted September 2-3, 2026, hashed every image with SHA-256 and compared duplicates within and across conversations. The findings raise questions about what OpenAI does with user-submitted image data in coding agent conversations.

Topics

AI securityChatGPT

Sources

Go deeper

This intelligence is sourced automatically from public sources across the web and synthesised by the Prefactor AI pipeline. Stories are reviewed before publication.