344
Productivity & Workflow355
Automation & Workflow225
Software Development251
Marketing & Growth192
AI Infrastructure & MLOps175
Writing & Content Creation203
Data & Analytics142
Photography & Imaging156
Design & Creative170
Customer Support132
Sales & Outreach125
Voice & Speech135
Education & Learning131
Operations & Admin87
Two mathematicians say OpenAI should prove its AI did not learn from their unpublished ideas or private chats while producing new math results.
In short: Some mathematicians want OpenAI to prove its AI models did not learn from unpublished research or private user chats.
A new dispute is growing over where OpenAI gets the information that helps its AI do advanced mathematics. Mathematician Andreas Thom wrote that OpenAI should be more transparent about whether user interactions and unpublished work end up in its training data.
Training data is the material used to teach an AI system, like giving a student lots of textbooks and worksheets to study. Thom said OpenAI showed a “detailed command” of specialized techniques in his field, and he wondered if earlier conversations he and colleagues had with ChatGPT could have played a role.
Thom’s concerns come after another mathematician, NYU professor Tristan Buckmaster, raised similar questions about whether OpenAI’s tools may have benefited from his work. OpenAI recently announced several math advances, including one related to “non-sofic groups,” which are infinite mathematical objects that cannot be closely approximated using finite ones (like trying to model an endless pattern using only a limited set of tiles).
Thom said OpenAI’s replies did not clearly address whether user conversations could be included indirectly in the large pools of data used to improve models. He argued that researchers cannot realistically check this themselves, and only OpenAI can provide the evidence.
The key issue is OpenAI’s wording about “de-identified” data, meaning data stripped of names or obvious personal details. Thom argues that removing a name does not remove the underlying idea, and he says it would be unethical if private, nonpublic research helped OpenAI publish results faster than the researchers who originated the work. OpenAI did not immediately respond to The Verge’s request for comment.
Source: The Verge AI