UNCOS

Apple Engineers Show How Flimsy AI ‘Reasoning’ Can Be

Apple Engineers Show How Flimsy AI ‘Reasoning’ Can Be

image via WIRED

October 15, 2024, 11:55 PM

  • A new study from Apple engineers shows that the mathematical "reasoning" displayed by advanced large language models can be extremely brittle and unreliable in the face of seemingly trivial changes to common benchmark problems.
  • The fragility highlighted in these new results helps support previous research suggesting that LLMs' use of probabilistic pattern matching is missing the formal understanding of underlying concepts needed for truly reliable mathematical reasoning capabilities.
  • The tested LLMs fared much worse, though, when the Apple researchers modified the GSM-Symbolic benchmark by adding "seemingly relevant but ultimately inconsequential statements" to the questions.

New research shows that the mathematical reasoning displayed by advanced large language models can be extremely brittle and unreliable in the face of seemingly trivial changes to common benchmark problems. The fragility highlighted in these new results helps support previous research suggesting that LLMs' use of probabilistic pattern matching is missing the formal understanding of underlying concepts needed for truly reliable mathematical reasoning capabilities.

Read original article

Entities Mentioned

Apple

Topics Covered

BusinessBusiness / Artificial Intelligence

Comments (0)

No comments yet.