Special Analysis
The strategic question in business applications of artificial intelligence is usually framed as a choice between systems. This week brought four events in which the choice itself was not the hardest part. The biggest challenge was evidence. In three cases, the organization making a high-stakes claim also controlled most of the evidence or assurances available to those expected to rely on that claim. The fourth case showed what a different approach can look like, one in which the evidence itself can be independently verified.
A $12.93 billion acquisition, a model launch, a federal audit and a mathematical proof seem to have little in common at first glance. Taken together, they describe a market in which high-stakes claims are thoroughly documented, but mostly by parties with an interest in the conclusions. For an organization that buys AI rather than builds it, the question is no longer just which model performs best, but what evidence it receives, who prepared it and what would have to be true to show that it is wrong.
On September 2, NVIDIA signed a definitive agreement to acquire Hugging Face for $12,930,300,000, announced it on September 3, and expects the transaction to close in the first half of 2027, subject to regulatory approval. Around $11.9 billion is earmarked for existing shareholders, while up to $1 billion is set aside for equity packages to retain employees who move to NVIDIA (NVIDIA blog, Jensen Huang, September 3, 2026); (The Register, September 3, 2026). According to NVIDIA's own figures, the platform hosts more than three million models and 500,000 datasets and is used by more than 18 million developers and researchers. In the announcement, Huang stated that the platform will remain open and that "NVIDIA computing infrastructure will not be required to develop or deploy through Hugging Face" (TechCrunch, September 3, 2026).
What has changed is not access to the platform, but its likely future ownership, and with it the assumptions organizations use to assess its long-term neutrality. Hugging Face is where much of model selection for business use actually takes place: weights, licenses, model cards and evaluation materials are compared there before anything reaches procurement documentation. If the transaction closes, that comparison layer will be owned by the largest supplier of the compute infrastructure those models run on. The promise of neutrality is a publicly stated intention of the acquiring company's CEO, not a contractual provision a buyer can invoke.
For organizations across Southeast Europe that adopted open-weight models to reduce dependence on a single vendor, the consequence is contractual, not emotional. An open-model strategy that relied on Hugging Face as a relatively neutral distribution layer must now account for the possibility that this layer will be controlled by the dominant supplier of AI compute. If implemented properly, such a strategy does not depend on whether the platform stays neutral: models, licenses and documentation can be copied and stored independently. The sensible response is to do exactly that and to document in advance how to proceed if the platform's terms of use change.
Vendor-controlled evidence
On September 3, OpenAI introduced GPT-6 Astra and announced that it is the company's first model to reach a critical capability level in cybersecurity under its Preparedness Framework. This means that, with the appropriate tools and access, it can find previously unknown security vulnerabilities and develop new ways to exploit them in well-protected systems, without detailed human guidance (OpenAI, Safety overview: GPT-6 Astra). Access is disabled by default in enterprise workspaces and must be enabled by an administrator (CSO Online, September 4, 2026).
The same system card states that monitorability has decreased compared to GPT-5.6 Sol, and that in situations deliberately designed to deceive it, the model may go undetected while strategically underperforming on evaluations (GPT-6 Astra system card). Referring to covert deliberate underperformance, OpenAI states that it "most likely would not be able to detect it" (Transformer, September 4, 2026). Disclosing this is to the company's credit, but it also means that the most significant capability assessment of the year was measured, documented and qualified by the very organization selling the product.
Sanchit Vir Gogia draws an even sharper conclusion. Astra's capabilities did not change between August 10, when OpenAI stated that the critical capability level could not be ruled out, and September 1, when it announced that the threshold had been reached: what changed was the testing, not the model. In his view, Astra is therefore the only frontier model whose cyber capabilities are actually known to enterprise users; models without such a designation, which already have access to enterprise accounts and credentials, have not been measured in the same way (CSO Online, September 4, 2026). The instinct to treat the flagged model as the risky one actually inverts the real situation.
Buyers in the region are entrusted with the risk management decision, but without the means to make it fully independently: most of the relevant evidence comes from the vendor, while independent testing can cover only part of the capability and risk profile. Although certain AI Act transparency obligations apply from August 2, 2026 (Mayer Brown, July 2026), these requirements should not be equated with independent technical confirmation of a model's capabilities or safety. The practical step is procedural: require that the vendor's threshold assessment, together with its stated limitations, becomes part of the contractual documentation, and treat a model that has not been measured as unmeasured, not as safe.
On September 3, Tesla began charging for driverless Cybercab rides in Austin, and on the same day the U.S. National Highway Traffic Safety Administration opened Audit Query AQ26002, which it made public on September 4. The agency is not examining how the vehicle drives. It is examining the process and technical data behind Tesla's own certification that the Cybercab meets federal safety standards, including the extent to which that certification relied on the position that certain standards do not apply to a vehicle without a steering wheel, pedals or mirrors (NHTSA statement, September 4, 2026). No recall or suspension of operations has been ordered, and the existing standards remain in force. Under the U.S. system, manufacturers self-certify compliance, while the regulator verifies it after the fact (TechCrunch, September 4, 2026).
This is not a story about a model, and that is exactly why it is useful. It is this week's clearest example of how self-certification works when a product is genuinely new: the verification is real, but it begins only after commercial deployment is already underway. European buyers should view it in light of their own regulatory framework. Many standalone high-risk AI systems under the EU AI Act rely largely on internal conformity assessment carried out by providers, although third-party assessment is mandatory for certain categories and product contexts. These Annex III obligations will apply from December 2, 2027, while the rules for high-risk AI systems embedded in regulated products will apply from August 2028, following the entry into force of the AI Omnibus on July 27, 2026 (European Commission, July 27, 2026); (Mayer Brown, July 2026). Serbia's draft law on artificial intelligence, presented in June 2026, is modeled on the same act and provides for conformity assessment, registration and continuous post-market monitoring of high-risk systems (NALED, June 11, 2026).
The lesson for an organization deploying a system is narrow and concrete. Protection does not come from the declaration of conformity itself, but from the documentation behind it: technical documentation, test data and a record of which requirements the vendor deemed not applicable and on what basis. That last item is exactly what NHTSA requested from Tesla, and it is the least likely to appear in vendor documentation unless the buyer explicitly asks for it.
Independently verifiable evidence
On September 4, Anthropic published the first complete, computer-verified proof of Fermat's Last Theorem, written in the Lean programming language. The proof was produced by Claude, largely autonomously, over 11 days. The formalization spans 13 million lines of Lean code and proves 29,500 auxiliary theorems along the way (Anthropic, September 4, 2026). Fermat's Last Theorem was proven by Andrew Wiles in 1995: this is a formalization based on an existing proof strategy, not a new theorem, and it was built on community infrastructure, including Mathlib and Kevin Buzzard's formalization project at Imperial College London.
The significance of this event is structural, not mathematical. Verifying the result depends far less on the author's reputation or interests, because the proof can be re-run against a precisely defined statement and a relatively small trusted software base. A certain level of trust is still required - in the formulation itself, the axioms and the Lean kernel - but it is trust of a different order than accepting a vendor's summary. It is almost the exact opposite of the previous three examples.
The limitation is just as clear: formal verification applies where claims can be formally stated, which excludes most of what a buyer wants to know about a model in production. Still, the direction is instructive for anyone drafting a procurement specification: evidence the recipient can re-check on their own is worth more than evidence they have to trust. Benchmarking tools that run in your own environment, logs your internal team can search and tests the organization itself controls are weaker than a Lean proof, but stronger than a vendor summary. Requiring them at the tender stage is far cheaper than introducing them after implementation.
Conclusion
To close, an analogy from the same week. On September 6, the Financial Times reported that UBS will require graduates and interns applying to its global banking and markets division for the 2027 intake to demonstrate how they can use AI to improve performance and work efficiency, with AI questions included in interviews (Financial Times, September 6, 2026). The bank has decided to verify young candidates' claims of AI proficiency during hiring, rather than assuming that competence or treating it as merely desirable.
Such a standard is not unreasonable. It is simply applied at the point in the chain where verification is cheapest. In the same week, three organizations responsible for far more significant platforms and systems produced or controlled much of the evidence on which initial decisions had to be based. No one is acting improperly: self-certification is the legal default in several of these regimes, and OpenAI voluntarily published the least favorable finding from its own documentation. The problem is not honesty. The problem is that many organizations buying AI have not yet developed a consistent practice of asking for evidence they can verify themselves.
For local organizations that will buy AI rather than build frontier models, that practice is all the leverage they have. They will not be able to audit a frontier AI lab or renegotiate a platform acquisition. But they can decide what goes into their contractual documentation. So the question to ask at your next vendor meeting is simple: which claims underlying your current AI systems can your team independently verify, and which do you accept only because the party making them said so?