Skip to content
Changelog

Reader fixes for PowerPoint, JSON and sitemaps

September 16, 20261 min read

Agno's PPTXReader, JSONReader and SitemapReader each got fixes for content they skipped or handled the wrong way.

PPTXReader reads text inside grouped shapes

PPTXReader used to read only the top level of each slide. It skipped text inside a group, and a slide whose text sat in one group read as (No text content). PPTXReader now reads the text in groups, including nested groups:

Before: Slide 1: Q3 review
Now:    Slide 1: Q3 review
        Revenue up 12%
        Churn down 3%

JSONReader.async_read() awaits async chunking

JSONReader.async_read() ran the synchronous reader in a thread. That reader called your strategy's chunk() even when the strategy provides an async achunk(). JSONReader.async_read() now parses the file off the event loop and awaits achunk().

SitemapReader keeps pages like myindex.html

SitemapReader removes index.html from page URLs. That step also cut index.html from the end of other filenames, so myindex.html became my. When a sitemap listed both pages, the reader skipped one and gave the other the wrong document ID. SitemapReader now removes only a real index.html or index.htm.

SitemapReader continues past a broken compressed sitemap

A truncated or corrupt .xml.gz file used to stop discovery in SitemapReader. The reader now tries the next sitemap. If SitemapReader can't read a child of a sitemap index, the other pages still load and the reader sets discovery_incomplete. That flag tells Agno Knowledge not to prune pages that discovery missed.

Learn more about readers, the PPTX reader and the JSON reader in the documentation.

Frequently asked questions

PPTXReader used to read only the top level of each slide, so it skipped text inside grouped shapes. Now it also reads text inside groups and nested groups.

Yes. JSONReader.async_read() now awaits your strategy's achunk() and parses the file off the event loop. It used to call the synchronous chunk().

SitemapReader moves on to the next sitemap. When a child of a sitemap index fails to read, the reader still loads the other pages and sets discovery_incomplete, so Agno Knowledge doesn't prune pages that discovery missed.

Shipped around the same time