← All articlesAI Lab

How FitFat AI Tests Wellness Answers Without Pretending AI Is a Doctor

The exact disclosure, prompting and fact-checking rules behind our model comparisons.

Published 2026-08-24 · Reviewed 2026-08-24 · 10 min read · FitFat AI Editorial Team
Editorial cover for How FitFat AI Tests Wellness Answers Without Pretending AI Is a Doctor

Ask several AI systems the same wellness question and the answers may look remarkably similar: a confident introduction, a numbered list and a reminder to consult a professional. The meaningful differences often sit underneath that polished surface. One model may identify an important exception. Another may invent a source. A third may give a technically correct answer that is almost impossible to use.

FitFat AI was created to document those differences honestly. We are not interested in publishing a disguised advertisement for whichever model produces the most exciting paragraph. We want to know what a real adult would receive, what can be verified and what still requires professional judgment.

We begin with a question, not a preferred winner

A comparison starts with a practical question that an adult might reasonably ask about movement, food, recovery, sleep or everyday wellness. The question should be specific enough to evaluate. “How can I be healthier?” is too broad; “How should a beginner divide walking and strength work across one week?” produces claims and recommendations that can be checked.

Before running the test, we write the evaluation criteria. This reduces the temptation to change the rules after seeing which answer we like.

Every model receives the same core prompt

We record the exact prompt, product, model name shown by the provider, test date and relevant settings. If web browsing is enabled for one system, that difference must be visible. If a product does not reveal the precise underlying model, we do not guess.

A new chat is used for each system so that unrelated conversation history does not influence the answer. We preserve the original output. Formatting may be shortened for presentation, but omissions must be marked and cannot change the meaning.

Follow-up questions are useful, particularly when an answer makes a questionable claim. They are recorded separately because a model that succeeds only after several corrections did not perform the same as one that handled the risk in its first response.

Fluent language is not evidence

We check important health claims against sources outside the AI response. Preferred references include public-health agencies, government health institutes, established professional medical organizations, systematic reviews and original peer-reviewed research.

A citation earns no credit merely because it looks academic. We check whether the page or paper exists, whether it says what the model claims and whether the evidence applies to the population in the question. A study in a small or very different group cannot automatically support a universal recommendation.

Safety and uncertainty matter

The strongest answer is not necessarily the most cautious or the longest. It should identify safety issues that materially affect the decision without turning every ordinary question into an emergency warning.

We look for reasonable boundaries: symptoms that require urgent help, conditions that may need individual modification, and situations where general guidance is not enough. We also examine whether the model distinguishes association from cause and whether it admits when evidence is limited.

A score belongs to one test

When we use a score, it applies to the documented prompt, date and settings. AI products change quickly, and output can vary between runs. A model that performs well on exercise programming may perform poorly on supplement evidence. We therefore avoid claims that one system is permanently “the smartest” or “the safest.”

Our comparison categories may include relevance, factual support, source integrity, safety context, uncertainty and practical usefulness. The article explains the scoring scale before presenting a result.

What we will never do

We will not invent model outputs, claim access we did not have, hide a sponsorship or rewrite a weak answer into a stronger one. We will not use AI consensus as proof that a health claim is true. Several models can repeat the same widespread error because they may draw on overlapping material.

The purpose is not to replace a clinician, dietitian, physiotherapist or qualified trainer. It is to help readers see what an AI answer can and cannot establish.

The reader should be able to audit us

A useful comparison leaves a trail: prompt, date, model disclosure, relevant output, evaluation criteria and independent verification sources. Readers should be able to disagree with our judgment while understanding how we reached it.

That transparency is the core of the FitFat AI identity. The interesting story is not that an AI produced a list. It is where the answers diverged, why the difference matters and what the best available evidence actually supports.

Editorial references


Reviewed under our editorial policy. See our AI comparison methodology. Please also read our medical disclaimer.