AI evaluation and benchmark collaboration offers
Category: Questions & membership Priority: P2 Area: ai-answers Reports: 3 Status: Open
Summary
Three technical outsiders independently offered to evaluate Qaf's AI quality and safety within a two-day window in early August: Islamify AI's co-founder invited Qaf onto the IslamicMMLU leaderboard (run by UK researchers with Dr. Waleed Kadous; their full test cost Islamify ~$200 and the dataset is deliberately not public), a Berkeley statistics PhD building safetyevidence.org offered to run safety/ethics/values evaluations or collaborate on internal evals, and a Spotify engineer specializing in agent/LLM evaluation and observability asked what eval metrics Qaf has and offered pro-bono help. Abdellatif's replies confirm the gap these offers address: Qaf has no fine-tuned model (it uses OpenAI models), no proper eval set yet ("mostly been doing manual testing and iteration"), and generation costs ~$0.20 per answer, making full-benchmark runs expensive. He deferred each offer by roughly two months but kept every door open — this cluster is effectively a queue of free expert help for when the eval set gets built.
What users say
Meris Cerić meris@islamify.ai — IslamicMMLU leaderboard invitation
- Date: 2026-08-08, Language: English, Thread:
19fe08ff3b94ef67, Replied: yes - Co-founder of Islamify AI invites Qaf to be evaluated on the IslamicMMLU benchmark (https://huggingface.co/spaces/islamicmmlu/leaderboard) so Islamic LLMs can be compared, and graciously predicts Qaf "might well take first place from us". He clarifies he doesn't run the leaderboard (UK brothers, with Dr. Waleed Kadous), that the full test cost Islamify ~$200, and that the dataset isn't released even as a subset to protect benchmark integrity. The thread then turns into friendly peer exchange: Abdellatif tried Islamify, gave UI feedback (font sizing, clickable citation pills), and Meris shared their read-aloud TTS implementation in detail.
"I recently came across Qaf and wanted to invite you to participate in the Islamic MMLU benchmark and have Qaf added to the leaderboard as well. It would be great to see more Islamic LLMs evaluated on the same benchmark and to make the results comparable across different models."
Anthony Ozerov ozerov@berkeley.edu — safety/ethics/values evaluation
- Date: 2026-08-06, Language: English, Thread:
19fd86165f7ca831, Replied: yes (2026-08-09) - A statistics PhD student at Berkeley building safetyevidence.org offers to run Qaf's model through his safety/ethics/values evaluations (if fine-tuned) or to collaborate on Qaf's internal evals. Abdellatif answered that Qaf has no fine-tuned model (currently using OpenAI's) and no proper eval set yet — "mostly been doing manual testing and iteration, but it's quite high on our priority list" — and said he'd love to chat once one exists. Anthony said to reach out whenever.
"I am a PhD student in statistics currently working on evaluating the safety, ethics, and values of AI models, and building safetyevidence.org. I saw qaf.ai and it is pretty cool! I think there are some ways we can work together."
Mohammad Salman salman8832@hotmail.com — LLM eval/observability, pro-bono offer
- Date: 2026-08-07, Language: English, Thread:
19fda24d86a170cb, Replied: yes (2026-08-13; original went to spam) - A Spotify engineer specializing in agent/LLM evaluation and observability asks what evaluation and observability metrics Qaf has in place and offers to support the team, later explicitly offering to work pro-bono "because I want a product like this to succeed". Abdellatif: "We don't do much evaluation at this time unfortunately. Let's talk in ~2 months if you're free."
"I wanted to ask about what evaluation and observability metrics have you put in place for this app? The reason I ask is because I specialize in Agent/LLM evaluation and observability at my day job at Spotify, and would love to support your team in any way possible!"
Details / repro
- Cost constraint on benchmarking: Qaf answers cost ~$0.20 per generation; Islamify's full IslamicMMLU run cost ~$200. IslamicMMLU dataset is not public (integrity), so no cheap subset run is possible.
- Referenced research: "IslamicMMLU: A Benchmark for Evaluating LLMs on Islamic Knowledge" (ResearchGate publication 403154641).
- Bonus intel from the Islamify thread (read-aloud/TTS, relevant to Qaf's planned feature): Gemini 3 Flash TTS (initially misstated as gpt-4o-mini-tts), routed through LiteLLM and an authenticated TTS endpoint; text cleaned and split by paragraph/punctuation; chunked generation with caching keyed on text+model+config; browser queues audio sections for continuous playback (not token-level streaming); pipeline is provider-agnostic.
- Current internal state per Abdellatif: OpenAI models, no eval set, manual testing only; eval set "quite high on our priority list". Follow-ups implicitly promised for ~October 2026 (Salman, Ozerov) — worth calendaring.
Threads
19fe08ff3b94ef67— Islamify co-founder: IslamicMMLU leaderboard invite + TTS implementation exchange19fd86165f7ca831— Berkeley PhD (safetyevidence.org): safety/ethics/values eval offer19fda24d86a170cb— Spotify engineer: eval/observability metrics question, pro-bono offer