We're Building AI for Rural America. It's Time to Measure Whether It Works.
By Emaliah Sawyer, Head of Product & Story, Center for Rural AI
The Colorado Thrives AI Fellowship
This summer, I had the opportunity to be a Colorado Thrives AI Fellow, hosted by the Center for Rural AI (CRAI) alongside a fellow CRAI colleague. The program began with training in Denver with Anthropic's Beneficial Deployment team and CodePath, followed by a ten-week build period with weekly CodePath mentorship meetings to complete our projects. After the training, we set off to develop systems that would help our organization advance its mission.
Before arriving for the in-person training in Denver, we held discovery conversations with our managers. That way, we could come into the training with a problem already identified, start digging into it, and think toward solutions we could build.
Setting Out to Build a Rural Benchmark
I identified a need for the Center for Rural AI to put numbers behind its mission, so I decided to build a rural benchmark evaluation: a way to measure how different LLMs represent rural communities. This was an important initiative. It would put quantitative evidence behind a problem that I and other rural community members were seeing, the same observation CRAI was founded on: biases and assumptions in AI that don't favor rural areas. Those assumptions compound the economic disadvantages rural communities already face with AI technology, and they start in the data and model training that power frontier LLMs.
I quickly found that this initiative was far too big for a ten-week summer fellowship. Done well, a full rural benchmark would look more like a multi-month team research project. That realization led me to collaborate with my colleague and co-fellow Holden Bronson, who was building a system called BRAIN.
What BRAIN Is
For the fellowship, Holden set out to build BRAIN. BRAIN is an MCP (Model Context Protocol) system that pulls on CRAI's rural AI data and plugs into your LLM, giving it access to additional rural context to enrich outputs for rural users. Right now, BRAIN is used internally at CRAI, and we're working to grow its data corpus. Holden's full story is worth reading on its own.
Instead of benchmarking every model against all of rural America, I partnered with Holden to evaluate BRAIN itself. That's what BRAIN Bench is: an evaluation system that measures how well BRAIN delivers on its goal of providing value to rural users. CRAI was investing in BRAIN based on a hypothesis of value, but it needed proof, along with quantitative next steps for improving how BRAIN works. BRAIN Bench tells us how BRAIN performs against raw LLMs, where the gaps are in the corpus, and whether changes to BRAIN make outputs better or worse.
How BRAIN Bench Works
Before we get into the mechanics of BRAIN Bench, let's start with a quick definition. A benchmark is a structured test. You give a system a set of questions, score the responses against a fixed rubric, and then use the scores to compare outputs, either against a baseline or against themselves over time. BRAIN Bench mirrors that structure, and it was built specifically to answer whether BRAIN makes a model's answers more useful for CRAI's rural AI work, and if so, by how much.
BRAIN Bench asks 101 sample prompts across six test subjects: Claude Opus, Sonnet, and Haiku; GPT-5.6 Sol and Terra; and BRAIN MCP + Sonnet. Each output is graded by an automated LLM judge running on Sonnet, using a rubric I built that scores four dimensions: accuracy, completeness, rural nuance, and honesty.
Accuracy asks whether what the answer claims is actually true. Completeness asks whether it covered everything the question asked for. Rural nuance asks whether the answer names something specific to a rural setting that actually changes the conclusion, rather than defaulting to a generic, metro-shaped answer. Honesty asks whether the model's confidence matched what it actually had evidence for. Each output is graded pass or fail, and a response only counts as a win if it passes all four. One failing dimension fails the whole answer.
BRAIN Bench runs as an automated system, with a structured workflow that guides each step of the evaluation and uses API calls to send prompts and receive outputs. As Holden keeps building out BRAIN's data and retrieval, the same harness can run again and show, in numbers, whether those changes moved the needle.
I also built a user interface so the whole team could easily look at the results, no code required. It's organized into five tabs:
- Dashboard: the headline view, with a leaderboard of every model, scores by rubric dimension, and score compared to response time.
- By Category: the same scoring, sliced by industry segment, such as Economy and Jobs, Agriculture, Healthcare, Schools, and Local Government.
- BRAIN Gaps: a BRAIN-only view showing where it fell short, including questions its corpus didn't cover at all, the weakest categories, and suggested improvements.
- Definitions: what's being measured, the grading rules, and who and what is being tested.
- Outputs: every model's full answer to every question, with the judge's pass or fail on each dimension, so anyone can check the grading for themselves.
There's also a toggle to compare models only on the questions where BRAIN had an answer in its corpus, so a “not in my corpus” reply doesn't skew the comparison.
Breaking the questions into industry segments turned out to be one of the most useful choices. By looking at how BRAIN scored in each segment, we can see which sectors of the corpus are well built out and which are the weakest. That tells us which industry data to focus on adding to the corpus next.
Along with finding ways to improve BRAIN through BRAIN Bench, we found that using an MCP like BRAIN to do these searches can save a lot of token burn. For organizations with limited AI budgets, that efficiency could be a real value add, and a real argument for building intentional data infrastructure to connect to LLMs.
What I Learned
More than anything, the fellowship made me a better AI builder, because I got to spend real time working with AI tools. A mentor of mine told me that time on tools is the key to building with new technology, and I couldn't agree more. You can spend all the time in the world learning what's out there and memorizing technical terms, but becoming an AI builder today means “rolling up your sleeves,” testing what's possible, and then reflecting on what went wrong and what went well. Technology is more accessible to more people than ever, and there are more resources than ever to help builders take on new and intimidating tasks, especially with AI.
BRAIN Bench was my time on tools. Whenever BRAIN Bench ran into a malfunction, I used AI as a thought partner and a tutor to help me fix each unexpected issue. Each fix taught me something I never could have learned just by reading about AI. Getting my hands “dirty” and actually building was critical.
Looking back, I didn't just build a benchmark this summer. I learned how to learn with AI: try something, see what breaks, figure out why, try again, and learn from it so I don't make the same mistake twice. That cycle is how I'll approach every project from here on out.
A Baby Step in the Right Direction
The wonderful thing about what Holden and I built is that these tools aren't just for the benefit of one organization. They were built with the purpose of helping other organizations and entire rural communities as well. Right now, we're working on how to roll BRAIN and BRAIN Bench into even bigger initiatives that will serve rural communities further, and in larger ways.
BRAIN Bench is open source. Check out the BRAIN Bench repo on GitHub to see how I built it and explore our other open-source resources. We love community engagement, and we welcome ideas and improvements from anyone who wants to help.
Learn more about the Center for Rural AI's work and the tools being built for rural communities.