1 points | by sameerhimati an hour ago
1 comments
I was picking a search API for a research agent and every comparison I could find was run by one of the vendors so I ran my own.
Most search benchmarks end up grading the answer the model produces, not the actual pages returned.
Open to feedback!
I was picking a search API for a research agent and every comparison I could find was run by one of the vendors so I ran my own.
Most search benchmarks end up grading the answer the model produces, not the actual pages returned.
Open to feedback!