Two newer ways to log food have shown up to replace typing: point a camera at your plate, or say what you ate out loud. Both are faster than searching a database and picking from a list of near matches.
Only one of them can actually tell what is in your food.
This post compares voice first tracking and photo based tracking on the things that decide whether a log is close to correct: what the method can observe, how portions get estimated, and how long each log takes in daily use.
How photo based tracking works
A photo app asks for a picture of your plate, then runs image recognition to guess what is on it. It identifies shapes and colors it has seen before: a chicken breast, a mound of rice, a green salad. From there it estimates a portion size and looks up calories and macros for each item.
This works reasonably well for simple, visually distinct plates. A single piece of grilled fish with steamed broccoli is easy to read from above.
What the camera can’t see
A photo is a single angle at a single moment, and that is the whole problem. The camera cannot see:
- Oil, butter, or dressing mixed into a dish rather than sitting on top
- Sauce underneath a protein or folded into a stew
- Ingredients under the top layer, like the filling inside a wrap or the meat under rice in a bowl
- How thick a portion actually is, since depth is hard to judge from one photo
- Anything you already ate before the picture was taken, like a spoon of peanut butter or a splash of milk in coffee
None of this is a flaw in any one app. It is a limit of the input. A photo shows a surface. Calories live in what you cannot see from that angle: the tablespoon of oil the food was cooked in, the sauce mixed through the noodles, the butter on the bread under the sandwich.
How voice first tracking works
Voice first tracking flips the input. Instead of showing the app your plate, you describe it: “a chicken wrap with hummus and a small side salad” or “two eggs, toast with butter, and a coffee with milk.” The app transcribes what you said, matches it to nutrition data, and estimates the portion from how you described the meal, a standard plate, a regular wrap, a 330ml drink.
Because you are describing the meal rather than showing a photo of it, you can say the parts a camera would miss. If there is a tablespoon of olive oil in the pan, you say so. If the sauce is mixed through the pasta rather than sitting on top, you say so. The estimate only improves because you supplied information a lens cannot capture.
This is the idea behind Logma: you say what you ate, it logs calories and macros in about 3 seconds. If you would rather not talk in a given moment, every voice prompt has a text fallback, and typing the same sentence gives the same result.
Speed compared, step by step
Speed is easier to compare directly, since it comes down to the number of steps between deciding to log and having a number on screen.
| Method | Typical steps | Rough time |
|---|---|---|
| Typing search (MyFitnessPal style) | Search term, scroll results, pick closest match, adjust weight, repeat per item | 30 to 90 seconds per meal |
| Photo based | Open camera, frame the plate, wait for recognition, review and correct the guess | 15 to 30 seconds per meal |
| Voice first | Say the meal in one sentence | About 3 seconds per item |
Typing based tracking is slow mainly because food databases have dozens of near duplicate entries for the same dish, and a mixed meal means repeating that search for every item on the plate. Photo based tracking removes the search step but adds a wait for recognition and often a correction step when the guess is off. Voice first tracking removes both: there is no database to search and no image to process, just a short sentence.
Accuracy compared
Speed only matters if the number you end up with is close to right. This is where the two newer methods diverge the most.
Photo based accuracy depends entirely on what is visible. A plain, simple plate photographs well. A stir fry, a curry, a sandwich with layers, or anything with a sauce mixed through it is much harder, because the app is guessing at ingredients and portions it cannot fully see. Regional and home cooked dishes are also harder to recognize from an image if the app was not trained on them.
Voice first accuracy depends on what you say. If you describe the meal with the detail that matters (the oil, the sauce, the rough portion), the estimate can account for it. Logma understands meals described in most major languages, with regional dishes recognized natively, so a home cooked dish from any cuisine is not a guess based on how it looks in a photo, it is based on what you said was in it.
Neither method is going to be perfect down to the gram. Portion estimation from a description, like portion estimation from a photo, is still an estimate. The difference is what information the estimate is allowed to use.
When each one makes sense
- Photo based tracking can be a reasonable fit if your meals are simple and visually clear most of the time, and you do not mind occasional corrections.
- Voice first tracking fits better if your meals are mixed, home cooked, or drawn from a cuisine that is hard to identify from a picture, or if you just want the fastest possible log.
- Typing based tracking still works if you already know exact brands and weights and do not mind the extra steps, since manual entry is exact when the data is exact. It is also the slowest of the three once a meal has more than one or two items, which is a big part of why so many people quit after a week, a pattern covered in more detail in why MyFitnessPal feels slow.
- If speed is the main reason you stop tracking, it is worth reading how logging a meal in 3 seconds changes whether tracking sticks at all, since the method that takes the least effort is usually the one you keep doing.
Whichever method you use, the starting point is the same: knowing roughly how many calories you need. The free calorie calculator uses your weight, height, age, and activity level to estimate that number before you worry about how you log.
FAQ
Is voice calorie tracking accurate?
It depends on how well you describe the meal, the same way photo accuracy depends on what the camera can see. Because you can mention oil, sauce, and portion size directly, voice tracking can account for details a photo cannot capture.
Can photo calorie apps see oil and sauce?
Not reliably. A camera captures the visible surface of a plate, so oil mixed into cooking, sauce folded through a dish, and layers under the top ingredient are usually invisible to image recognition and get left out of the estimate.
Which is faster, voice or photo logging?
Voice is typically faster, since there is no photo to frame or wait on. A short spoken sentence can log a meal in about 3 seconds, compared to 15 to 30 seconds for a photo, including the review step.
Do I have to talk out loud to use a voice tracker?
No. Every voice prompt has a text fallback in apps built around this method, so typing the same description gives the same result if you are somewhere you cannot speak.