AI

The takeaways from Stanford’s 386-page report on the state of AI

Comment

Image Credits: Getty Images

Writing a report on the state of AI must feel a lot like building on shifting sands: By the time you hit publish, the whole industry has changed under your feet. But there are still important trends and takeaways in Stanford’s 386-page bid to summarize this complex and fast-moving domain.

The AI Index, from the Institute for Human-Centered Artificial Intelligence, worked with experts from academia and private industry to collect information and predictions on the matter. As a yearly effort (and by the size of it, you can bet they’re already hard at work laying out the next one), this may not be the freshest take on AI, but these periodic broad surveys are important to keep one’s finger on the pulse of industry.

This year’s report includes “new analysis on foundation models, including their geopolitics and training costs, the environmental impact of AI systems, K-12 AI education, and public opinion trends in AI,” plus a look at policy in a hundred new countries.

Let us just bullet the highest-level takeaways:

  • AI development has flipped over the last decade from academia-led to industry-led, by a large margin, and this shows no sign of changing.
  • It’s becoming difficult to test models on traditional benchmarks and a new paradigm may be needed here.
  • The energy footprint of AI training and use is becoming considerable, but we have yet to see how it may add efficiencies elsewhere.
  • The number of “AI incidents and controversies” has increased by a factor of 26 since 2012, which actually seems a bit low.
  • AI-related skills and job postings are increasing, but not as fast as you’d think.
  • Policymakers, however, are falling over themselves trying to write a definitive AI bill, a fool’s errand if there ever was one.
  • Investment has temporarily stalled, but that’s after an astronomic increase over the last decade.
  • More than 70% of Chinese, Saudi, and Indian respondents felt AI had more benefits than drawbacks. Americans? 35%.

But the report goes into detail on many topics and subtopics and is quite readable and nontechnical. Only the dedicated will read all 386 pages of analysis, but really, just about any motivated body could.

Let’s look at Chapter 3, Technical AI Ethics, in a bit more detail.

Bias and toxicity are hard to reduce to metrics, but as far as we can define and test models for these things, it is clear that “unfiltered” models are much, much easier to lead into problematic territory. Instruction tuning, which is to say adding a layer of extra prep (such as a hidden prompt) or passing the model’s output through a second mediator model, is effective at improving this issue, but it’s far from perfect.

The increase in “AI incidents and controversies” alluded to in the bullets is best illustrated by this diagram:

Image Credits: Stanford HAI

As you can see, the trend is upward and these numbers came before the mainstream adoption of ChatGPT and other large language models, not to mention the vast improvement in image generators. You can be sure that the 26x increase is just the start.

Making models more fair or unbiased in one way may have unexpected consequences in other metrics, as this diagram shows:

Image Credits: Stanford HAI

As the report notes, “Language models which perform better on certain fairness benchmarks tend to have worse gender bias.” Why? It’s hard to say, but it just goes to show that optimization is not as simple as everyone hopes. There is no simple solution to improving these large models, partly because we don’t really understand how they work.

Fact-checking is one of those domains that sounds like a natural fit for AI: Having indexed much of the web, it can evaluate statements and return a confidence that they are supported by truthful sources, and so on. This is very far from the case. AI actually is particularly bad at evaluating factuality and the risk is not so much that they will be unreliable checkers, but that they will themselves become powerful sources of convincing misinformation. A number of studies and datasets have been created to test and improve AI fact-checking, but so far we’re still more or less where we started.

Fortunately, there’s a large uptick in interest here, for the obvious reason that if people feel they can’t trust AI, the whole industry is set back. There’s been a tremendous increase in submissions at the ACM Conference on Fairness, Accountability, and Transparency, and at NeurIPS issues like fairness, privacy, and interpretability are getting more attention and stage time.

These highlights of highlights leave a lot of detail on the table. The HAI team has done a great job of organizing the content, however, and after perusing the high-level stuff here, you can download the full paper and get deeper into any topic that piques your interest.

More TechCrunch

For Mark Zuckerberg’s fortieth birthday, his wife got him a photoshoot. Zuckerberg gives the camera a sly smile as he sits amid a carefully crafted recreation of his childhood bedroom.…

Mark Zuckerberg’s makeover: midlife crisis or carefully crafted rebrand?

Strava announced a slew of features, including AI to weed out leaderboard cheats, a new ‘family’ subscription plan, dark mode and more.

Strava taps AI to weed out leaderboard cheats; unveils ‘family’ plan, dark mode and more

We all fall down sometimes. Astronauts are no exception. You need to be in peak physical condition for space travel, but bulky space suits and lower gravity levels can be…

Astronauts fall over. Robotic limbs can help them back up.

Microsoft will launch its custom Cobalt 100 chips to customers as a public preview at its Build conference next week, TechCrunch has learned. In an analyst briefing ahead of Build,…

Microsoft’s custom Cobalt chips will come to Azure next week

What a wild week for transportation news! It was a smorgasbord of news that seemed to touch every sector and theme in transportation.

Tesla keeps cutting jobs and the feds probe Waymo

Sony Music Group has sent letters to more than 700 tech companies and music streaming services to warn them not to use its music to train AI without explicit permission.…

Sony Music warns tech companies over ‘unauthorized’ use of its content to train AI

Winston Chi, Butter’s founder and CEO, told TechCrunch that “most parties, including our investors and us, are making money” from the exit.

GrubMarket buys Butter to give its food distribution tech an AI boost

The investor lawsuit is related to Bolt securing a $30 million personal loan to Ryan Breslow, which was later defaulted on.

Bolt founder Ryan Beslow wants to settle an investor lawsuit by returning $37 million worth of shares

Meta, the parent company of Facebook, launched an enterprise version of the prominent social network in 2015. It always seemed like a stretch for a company built on a consumer…

With the end of Workplace, it’s fair to wonder if Meta was ever serious about the enterprise

X, formerly Twitter, turned TweetDeck into X Pro and pushed it behind a paywall. But there is a new column-based social media tool in the town, and it’s from Instagram…

Meta Threads is testing pinned columns on the web, similar to the old TweetDeck

As part of 2024’s Accessibility Awareness Day, Google is showing off some updates to Android that should be useful to folks with mobility or vision impairments. Project Gameface allows gamers…

Google expands hands-free and eyes-free interfaces on Android

A hacker listed the data allegedly breached from Samco on a known cybercrime forum.

Hacker claims theft of India’s Samco account data

A top European privacy watchdog is investigating following the recent breaches of Dell customers’ personal information, TechCrunch has learned.  Ireland’s Data Protection Commission (DPC) deputy commissioner Graham Doyle confirmed to…

Ireland privacy watchdog confirms Dell data breach investigation

Ampere and Qualcomm aren’t the most obvious of partners. Both, after all, offer Arm-based chips for running data center servers (though Qualcomm’s largest market remains mobile). But as the two…

Ampere teams up with Qualcomm to launch an Arm-based AI server

At Google’s I/O developer conference, the company made its case to developers — and to some extent, consumers — why its bets on AI are ahead of rivals. At the…

Google I/O was an AI evolution, not a revolution

TechCrunch Disrupt has always been the ultimate convergence point for all things startup and tech. In the bustling world of innovation, it serves as the “big top” tent, where entrepreneurs,…

Meet the Magnificent Six: A tour of the stages at Disrupt 2024

There’s apparently a lot of demand for an on-demand handyperson. Khosla Ventures and Pear VC have just tripled down on their investment in Honey Homes, which offers up a dedicated…

Khosla Ventures, Pear VC triple down on Honey Homes, a smart way to hire a handyman

TikTok is testing the ability for users to upload 60-minute videos, the company confirmed to TechCrunch on Thursday. The feature is available to a limited group of users in select…

TikTok tests 60-minute video uploads as it continues to take on YouTube

Flock Safety is a multibillion-dollar startup that’s got eyes everywhere. As of Wednesday, with the company’s new Solar Condor cameras, those eyes are solar-powered and using wireless 5G networks to…

Flock Safety’s solar-powered cameras could make surveillance more widespread

Since he was very young, Bar Mor knew that he would inevitably do something with real estate. His family was involved in all types of real estate projects, from ground-up…

Agora raises $34M Series B to keep building the Carta for real estate

Poshmark, the social commerce site that lets people buy and sell new and used items to each other, launched a paid marketing tool on Thursday, giving sellers the ability to…

Poshmark’s ‘Promoted Closet’ tool lets sellers boost all their listings at once

Google is launching a Gemini add-on for educational institutes through Google Workspace.

Google adds Gemini to its Education suite

More money for the generative AI boom: Y Combinator-backed developer infrastructure startup Recall.ai announced Thursday it has raised a $10 million Series A funding round, bringing its total raised to over…

YC-backed Recall.ai gets $10M Series A to help companies use virtual meeting data

Engineers Adam Keating and Jeremy Andrews were tired of using spreadsheets and screenshots to collab with teammates — so they launched a startup, CoLab, to build a better way. The…

CoLab’s collaborative tools for engineers line up $21M in new funding

Reddit announced on Wednesday that it is reintroducing its awards system after shutting down the program last year. The company said that most of the mechanisms related to awards will…

Reddit reintroduces its awards system

Sigma Computing, a startup building a range of data analytics and business intelligence tools, has raised $200 million in a fresh VC round.

Sigma is building a suite of collaborative data analytics tools

European Union enforcers of the bloc’s online governance regime, the Digital Services Act (DSA), said Thursday they’re closely monitoring disinformation campaigns on the Elon Musk-owned social network X (formerly Twitter)…

EU ‘closely’ monitoring X in wake of Fico shooting as DSA disinfo probe rumbles on

Wind is the largest source of renewable energy in the U.S., according to the U.S. Energy Information Administration, but wind farms come with an environmental cost as wind turbines can…

Spoor uses AI to save birds from wind turbines

The key to taking on legacy players in the financial technology industry may be to go where they have not gone before. That’s what Chicago-based Aeropay is doing. The provider…

Cannabis industry and gaming payments startup Aeropay is now offering an alternative to Mastercard and Visa

Facebook and Instagram are under formal investigation in the European Union over child protection concerns, the Commission announced Thursday. The proceedings follow a raft of requests for information to parent…

EU opens child safety probes of Facebook and Instagram, citing addictive design concerns