# Reader fixes for PowerPoint, JSON and sitemaps

> Agno's PPTXReader now reads text inside grouped shapes, JSONReader async reads use your strategy's async chunking, and SitemapReader keeps pages like myindex.html and continues past a broken compressed sitemap.

- Published: 2026-09-16
- Author: Yash Raj Pandey
- Categories: Changelog
- Canonical: https://www.agno.com/articles/reader-fixes-for-powerpoint-json-and-sitemaps
- Markdown: https://www.agno.com/articles/reader-fixes-for-powerpoint-json-and-sitemaps.md

Agno's `PPTXReader`, `JSONReader` and `SitemapReader` each got fixes for content they skipped or handled the wrong way.

### `PPTXReader` reads text inside grouped shapes

`PPTXReader` used to read only the top level of each slide. It skipped text inside a group, and a slide whose text sat in one group read as `(No text content)`. `PPTXReader` now reads the text in groups, including nested groups:

```text
Before: Slide 1: Q3 review
Now:    Slide 1: Q3 review
        Revenue up 12%
        Churn down 3%
```

### `JSONReader.async_read()` awaits async chunking

`JSONReader.async_read()` ran the synchronous reader in a thread. That reader called your strategy's `chunk()` even when the strategy provides an async `achunk()`. `JSONReader.async_read()` now parses the file off the event loop and awaits `achunk()`.

### `SitemapReader` keeps pages like `myindex.html`

`SitemapReader` removes `index.html` from page URLs. That step also cut `index.html` from the end of other filenames, so `myindex.html` became `my`. When a sitemap listed both pages, the reader skipped one and gave the other the wrong document ID. `SitemapReader` now removes only a real `index.html` or `index.htm`.

### `SitemapReader` continues past a broken compressed sitemap

A truncated or corrupt `.xml.gz` file used to stop discovery in `SitemapReader`. The reader now tries the next sitemap. If `SitemapReader` can't read a child of a sitemap index, the other pages still load and the reader sets `discovery_incomplete`. That flag tells Agno Knowledge not to prune pages that discovery missed.

Learn more about [readers](https://docs.agno.com/knowledge/concepts/readers/overview), the [PPTX reader](https://docs.agno.com/reference/knowledge/reader/pptx) and the [JSON reader](https://docs.agno.com/knowledge/concepts/readers/json-reader) in the documentation.
