AI evaluation and benchmark collaboration offers

Category: Questions & membership Priority: P2 Area: ai-answers Reports: 3 Status: Open

Summary

Three technical outsiders independently offered to evaluate Qaf's AI quality and safety within a two-day window in early August: Islamify AI's co-founder invited Qaf onto the IslamicMMLU leaderboard (run by UK researchers with Dr. Waleed Kadous; their full test cost Islamify ~$200 and the dataset is deliberately not public), a Berkeley statistics PhD building safetyevidence.org offered to run safety/ethics/values evaluations or collaborate on internal evals, and a Spotify engineer specializing in agent/LLM evaluation and observability asked what eval metrics Qaf has and offered pro-bono help. Abdellatif's replies confirm the gap these offers address: Qaf has no fine-tuned model (it uses OpenAI models), no proper eval set yet ("mostly been doing manual testing and iteration"), and generation costs ~$0.20 per answer, making full-benchmark runs expensive. He deferred each offer by roughly two months but kept every door open — this cluster is effectively a queue of free expert help for when the eval set gets built.

What users say

Meris Cerić meris@islamify.ai — IslamicMMLU leaderboard invitation

"I recently came across Qaf and wanted to invite you to participate in the Islamic MMLU benchmark and have Qaf added to the leaderboard as well. It would be great to see more Islamic LLMs evaluated on the same benchmark and to make the results comparable across different models."

Anthony Ozerov ozerov@berkeley.edu — safety/ethics/values evaluation

"I am a PhD student in statistics currently working on evaluating the safety, ethics, and values of AI models, and building safetyevidence.org. I saw qaf.ai and it is pretty cool! I think there are some ways we can work together."

Mohammad Salman salman8832@hotmail.com — LLM eval/observability, pro-bono offer

"I wanted to ask about what evaluation and observability metrics have you put in place for this app? The reason I ask is because I specialize in Agent/LLM evaluation and observability at my day job at Spotify, and would love to support your team in any way possible!"

Details / repro

Threads