Skip to content
Changelog

Reader fixes for JSON, CSV and Tavily

September 23, 20261 min read

Every output below is real, from Agno's JSONReader, CSVReader, FieldLabeledCSVReader and TavilyReader on the previous release and on this one.

JSONReader keeps a scalar JSON root as one document

A JSON file can hold a single string, number, boolean or null at its root. JSONReader used to wrap only an object in a list before it built documents. A string root became one document per character, an empty string produced no documents, and a number, boolean or null raised TypeError. JSONReader now turns every root that is not an array into one document:

File                          Before                                  Now
"Agno docs home page"         19 docs: '"A"', '"g"', '"n"', '"o"'...  1 doc: '"Agno docs home page"'
42                            TypeError 'int' object is not iterable  1 doc: '42'
true                          TypeError 'bool' object is not iterable 1 doc: 'true'
null                          TypeError 'NoneType' object is not ...  1 doc: 'null'
""                            0 docs                                  1 doc: '""'

Objects and arrays read the same as before.

CSVReader reads text streams

CSVReader.read() and CSVReader.async_read() called .decode() on whatever the file object returned. A StringIO or a file opened in text mode returns a str, so the call failed. The reader caught the error, logged it and returned an empty list. CSVReader now decodes only bytes and passes text through:

Before: ERROR Error reading <_io.StringIO ...>: 'str' object has no attribute 'decode'
        read(StringIO) -> 0 docs
Now:    read(StringIO) -> 1 docs ['name, role\nAda, engineer\nLin, designer']

Both CSV readers reject a negative page_size

CSVReader.async_read() and FieldLabeledCSVReader.async_read() split a file with more than 10 rows into pages of page_size rows. A negative page_size produced no pages, so the reader returned no documents and logged nothing. Both readers now check page_size before they touch the file and raise ValueError, which leaves your stream where it was:

Before: CSVReader.async_read(page_size=-1)             -> 0 docs
        FieldLabeledCSVReader.async_read(page_size=-1) -> 0 docs
Now:    CSVReader.async_read(page_size=-1)             -> ValueError page_size cannot be a negative value.
        FieldLabeledCSVReader.async_read(page_size=-1) -> ValueError page_size cannot be a negative value.

A page_size of 0 on a file with more than 10 rows still logs range() arg 3 must not be zero and returns no documents, so pass a positive value.

TavilyReader sends its extract options

TavilyReader takes extract_depth and extract_format, but its request to Tavily Extract sent the depth under a depth key and left out the format. Tavily never received either setting. TavilyReader now sends extract_depth and format on both read() and async_read(). Values in params still override them. The request for TavilyReader(extract_depth="advanced", extract_format="text"):

Before: {'urls': ['https://docs.agno.com'], 'depth': 'advanced'}
Now:    {'urls': ['https://docs.agno.com'], 'extract_depth': 'advanced', 'format': 'text'}

See the cookbook, and learn more about readers, the CSV reader and the JSON reader in the documentation.

Frequently asked questions

JSONReader used to wrap only JSON objects in a list, so it iterated a string root character by character and raised TypeError on a number, boolean or null. Now any root other than an array becomes a single document.

CSVReader called .decode() on whatever the stream returned, which failed on a text stream with 'str' object has no attribute 'decode'. The reader then caught that error and handed back an empty list. CSVReader.read() and CSVReader.async_read() now decode only bytes.

CSVReader.async_read() and FieldLabeledCSVReader.async_read() now raise ValueError: page_size cannot be a negative value. before they read the file. On a file with more than 10 rows, a negative page_size used to return no documents with no error.

Yes. TavilyReader now sends extract_depth and format to Tavily Extract. It used to send the depth under a depth key and leave out the format.

Shipped around the same time