If it wasnt bad enough that Moores Law improvements in the density and cost of transistors is slowing. At the same time, the cost of designing chips and of the factories that are used to etch them is also on the rise. Any savings on any of these fronts will be most welcome to keep IT innovation leaping ahead.
One of the promising frontiers of research right now in chip design is using machine learning techniques to actually help with some of the tasks in the design process. We will be discussing this at our upcoming The Next AI Platform event in San Jose on March 10 with Elias Fallon, engineering director at Cadence Design Systems. (You can see the full agenda and register to attend at this link; we hope to see you there.) The use of machine learning in chip design was also one of the topics that Jeff Dean, a senior fellow in the Research Group at Google who has helped invent many of the hyperscalers key technologies, talked about in his keynote address at this weeks 2020 International Solid State Circuits Conference in San Francisco.
Google, as it turns out, has more than a passing interest in compute engines, being one of the large consumers of CPUs and GPUs in the world and also the designer of TPUs spanning from the edge to the datacenter for doing both machine learning inference and training. So this is not just an academic exercise for the search engine giant and public cloud contender particularly if it intends to keep advancing its TPU roadmap and if it decides, like rival Amazon Web Services, to start designing its own custom Arm server chips or decides to do custom Arm chips for its phones and other consumer devices.
With a certain amount of serendipity, some of the work that Google has been doing to run machine learning models across large numbers of different types of compute engines is feeding back into the work that it is doing to automate some of the placement and routing of IP blocks on an ASIC. (It is wonderful when an idea is fractal like that. . . .)
While the pod of TPUv3 systems that Google showed off back in May 2018 can mesh together 1,024 of the tensor processors (which had twice as many cores and about a 15 percent clock speed boost as far as we can tell) to deliver 106 petaflops of aggregate 16-bit half precision multiplication performance (with 32-bit accumulation) using Googles own and very clever bfloat16 data format. Those TPUv3 chips are all cross-coupled using a 3232 toroidal mesh so they can share data, and each TPUv3 core has its own bank of HBM2 memory. This TPUv3 pod is a huge aggregation of compute, which can do either machine learning training or inference, but it is not necessarily as large as Google needs to build. (We will be talking about Deans comments on the future of AI hardware and models in a separate story.)
Suffice it to say, Google is hedging with hybrid architectures that mix CPUs and GPUs and perhaps someday other accelerators for reinforcement learning workloads, and hence the research that Dean and his peers at Google have been involved in that are also being brought to bear on ASIC design.
One of the trends is that models are getting bigger, explains Dean. So the entire model doesnt necessarily fit on a single chip. If you have essentially large models, then model parallelism dividing the model up across multiple chips is important, and getting good performance by giving it a bunch of compute devices is non-trivial and it is not obvious how to do that effectively.
It is not as simple as taking the Message Passing Interface (MPI) that is used to dispatch work on massively parallel supercomputers and hacking it onto a machine learning framework like TensorFlow because of the heterogeneous nature of AI iron. But that might have been an interesting way to spread machine learning training workloads over a lot of compute elements, and some have done this. Google, like other hyperscalers, tends to build its own frameworks and protocols and datastores, informed by other technologies, of course.
Device placement meaning, putting the right neural network (or portion of the code that embodies it) on the right device at the right time for maximum throughput in the overall application is particularly important as neural network models get bigger than the memory space and the compute oomph of a single CPU, GPU, or TPU. And the problem is getting worse faster than the frameworks and hardware can keep up. Take a look:
The number of parameters just keeps growing and the number of devices being used in parallel also keeps growing. In fact, getting 128 GPUs or 128 TPUv3 processors (which is how you get the 512 cores in the chart above) to work in concert is quite an accomplishment, and is on par with the best that supercomputers could do back in the era before loosely coupled, massively parallel supercomputers using MPI took over and federated NUMA servers with actual shared memory were the norm in HPC more than two decades ago. As more and more devices are going to be lashed together in some fashion to handle these models, Google has been experimenting with using reinforcement learning (RL), a special subset of machine learning, to figure out where to best run neural network models at any given time as model ensembles are running on a collection of CPUs and GPUs. In this case, an initial policy is set for dispatching neural network models for processing, and the results are then fed back into the model for further adaptation, moving it toward more and more efficient running of those models.
In 2017, Google trained an RL model to do this work (you can see the paper here) and here is what the resulting placement looked like for the encoder and decoder, and the RL model to place the work on the two CPUs and four GPUs in the system under test ended up with 19.3 percent lower runtime for the training runs compared to the manually placed neural networks done by a human expert. Dean added that this RL-based placement of neural network work on the compute engines does kind of non-intuitive things to achieve that result, which is what seems to be the case with a lot of machine learning applications that, nonetheless, work as well or better than humans doing the same tasks. The issue is that it cant take a lot of RL compute oomph to place the work on the devices to run the neural networks that are being trained themselves. In 2018, Google did research to show how to scale computational graphs to over 80,000 operations (nodes), and last year, Google created what it calls a generalized device placement scheme for dataflow graphs with over 50,000 operations (nodes).
Then we start to think about using this instead of using it to place software computation on different computational devices, we started to think about it for could we use this to do placement and routing in ASIC chip design because the problems, if you squint at them, sort of look similar, says Dean. Reinforcement learning works really well for hard problems with clear rules like Chess or Go, and essentially we started asking ourselves: Can we get a reinforcement learning model to successfully play the game of ASIC chip layout?
There are a couple of challenges to doing this, according to Dean. For one thing, chess and Go both have a single objective, which is to win the game and not lose the game. (They are two sides of the same coin.) With the placement of IP blocks on an ASIC and the routing between them, there is not a simple win or lose and there are many objectives that you care about, such as area, timing, congestion, design rules, and so on. Even more daunting is the fact that the number of potential states that have to be managed by the neural network model for IP block placement is enormous, as this chart below shows:
Finally, the true reward function that drives the placement of IP blocks, which runs in EDA tools, takes many hours to run.
And so we have an architecture Im not going to get a lot of detail but essentially it tries to take a bunch of things that make up a chip design and then try to place them on the wafer, explains Dean, and he showed off some results of placing IP blocks on a low-powered machine learning accelerator chip (we presume this is the edge TPU that Google has created for its smartphones), with some areas intentionally blurred to keep us from learning the details of that chip. We have had a team of human experts places this IP block and they had a couple of proxy reward functions that are very cheap for us to evaluate; we evaluated them in two seconds instead of hours, which is really important because reinforcement learning is one where you iterate many times. So we have a machine learning-based placement system, and what you can see is that it sort of spreads out the logic a bit more rather than having it in quite such a rectangular area, and that has enabled it to get improvements in both congestion and wire length. And we have got comparable or superhuman results on all the different IP blocks that we have tried so far.
Note: I am not sure we want to call AI algorithms superhuman. At least if you dont want to have it banned.
Anyway, here is how that low-powered machine learning accelerator for the RL network versus people doing the IP block placement:
And here is a table that shows the difference between doing the placing and routing by hand and automating it with machine learning:
And finally, here is how the IP block on the TPU chip was handled by the RL network compared to the humans:
Look at how organic these AI-created IP blocks look compared to the Cartesian ones designed by humans. Fascinating.
Now having done this, Google then asked this question: Can we train a general agent that is quickly effective at placing a new design that it has never seen before? Which is precisely the point when you are making a new chip. So Google tested this generalized model against four different IP blocks from the TPU architecture and then also on the Ariane RISC-V processor architecture. This data pits people working with commercial tools and various levels tuning on the model:
And here is some more data on the placement and routing done on the Ariane RISC-V chips:
You can see that experience on other designs actually improves the results significantly, so essentially in twelve hours you can get the darkest blue bar, Dean says, referring to the first chart above, and then continues with the second chart above. And this graph showing the wireline costs where we see if you train from scratch, it actually takes the system a little while before it sort of makes some breakthrough insight and was able to significantly drop the wiring cost, where the pretrained policy has some general intuitions about chip design from seeing other designs and people that get to that level very quickly.
Just like we do ensembles of simulations to do better weather forecasting, Dean says that this kind of AI-juiced placement and routing of IP block sin chip design could be used to quickly generate many different layouts, with different tradeoffs. And in the event that some feature needs to be added, the AI-juiced chip design game could re-do a layout quickly, not taking months to do it.
And most importantly, this automated design assistance could radically drop the cost of creating new chips. These costs are going up exponentially, and data we have seen (thanks to IT industry luminary and Arista Networks chairman and chief technology officer Andy Bechtolsheim), an advanced chip design using 16 nanometer processes cost an average of $106.3 million, shifting to 10 nanometers pushed that up to $174.4 million, and the move to 7 nanometers costs $297.8 million, with projections for 5 nanometer chips to be on the order of $542.2 million. Nearly half of that cost has been and continues to be for software. So we know where to target some of those costs, and machine learning can help.
The question is will the chip design software makers embed AI and foster an explosion in chip designs that can be truly called Cambrian, and then make it up in volume like the rest of us have to do in our work? It will be interesting to see what happens here, and how research like that being done by Google will help.
See the original post here:
Google Teaches AI To Play The Game Of Chip Design - The Next Platform
- Cheating in Online Chess (Part 1): Suspicions of Engine Assistance - Chess.com - May 11th, 2024 [May 11th, 2024]
- Cheating in Online Chess (Part II): The Analysis of Engine Use - Chess.com - May 11th, 2024 [May 11th, 2024]
- The Silicon Gambit: How AI is Reshaping the World's Oldest Game - Chess.com - April 24th, 2024 [April 24th, 2024]
- Gukesh wins Candidates: The boy raised without chess engines wholl challenge Ding Liren at World Championships - The Indian Express - April 24th, 2024 [April 24th, 2024]
- Stars of the future shine in chess's ancestral homeland - Washington Times - September 19th, 2023 [September 19th, 2023]
- The 15 Best Episodes of Cowboy Bebop - MovieWeb - September 19th, 2023 [September 19th, 2023]
- Charge of the knight brigade: Indian teens storm global chess - IndiaTimes - August 20th, 2023 [August 20th, 2023]
- Knowing when to insist - ChessBase - August 20th, 2023 [August 20th, 2023]
- World Cup: Pragg and Salimova win tiebreakers - ChessBase - August 20th, 2023 [August 20th, 2023]
- What do F-16 and MiG-29 fighter jets do? - Times of Oman - August 20th, 2023 [August 20th, 2023]
- Xbox game releases August 21 to 27 - TrueAchievements - August 20th, 2023 [August 20th, 2023]
- Go! Guide Aug. 17 - The Republic - August 20th, 2023 [August 20th, 2023]
- MinStrength: An Alternative to Performance Rating - ChessBase - June 2nd, 2023 [June 2nd, 2023]
- Mittens (chess engine) - Wikipedia - January 31st, 2023 [January 31st, 2023]
- AlphaZero - Chess Engines - Chess.com - December 28th, 2022 [December 28th, 2022]
- 2022 U.S. Chess Championships, Round 3: Earning Respect! | US Chess.org - uschess.org - October 13th, 2022 [October 13th, 2022]
- Go! Guide Oct. 13 - The Republic - October 13th, 2022 [October 13th, 2022]
- Events, sales and more things happening Downriver The News Herald - Southgate News Herald - October 13th, 2022 [October 13th, 2022]
- Chess cheating drama: What are the different ways to cheat in chess? - The Indian Express - September 11th, 2022 [September 11th, 2022]
- Formula 1 2022: How to Watch the Italian Grand Prix Today - CNET - September 11th, 2022 [September 11th, 2022]
- The Machines That Made 500 Years of Circumnavigation Possible - Popular Mechanics - September 11th, 2022 [September 11th, 2022]
- Formula 1 2022: How to Watch the Belgian Grand Prix Today - CNET - August 29th, 2022 [August 29th, 2022]
- Kids want to grow, learn; are we planting seeds of knowledge? - Las Cruces Sun-News - August 29th, 2022 [August 29th, 2022]
- New: 3.h4 against the Kings Indian and Grnfeld - ChessBase India - August 25th, 2022 [August 25th, 2022]
- A bright chess champ emerges from Thiruvallur - The New Indian Express - August 25th, 2022 [August 25th, 2022]
- Interviewing The Coach Of Olympiad Sensation Gukesh - Chess.com - August 25th, 2022 [August 25th, 2022]
- Virtual Psychiatry is Here to Stay - Psychiatric Times - August 25th, 2022 [August 25th, 2022]
- Whatever Happened to the Transhumanists? - Gizmodo - August 2nd, 2022 [August 2nd, 2022]
- Beyond Carlsen: the devaluation of the World Chess Championship - TheArticle - July 31st, 2022 [July 31st, 2022]
- Go! Guide July 21 - The Republic - July 27th, 2022 [July 27th, 2022]
- Chennai Chess Olympiad and AI - Analytics India Magazine - June 24th, 2022 [June 24th, 2022]
- Go! Guide July 23 - The Republic - June 24th, 2022 [June 24th, 2022]
- Was Basman right? Iconoclasm, ridicule and chess - TheArticle - June 20th, 2022 [June 20th, 2022]
- Formula 1 Canadian Grand Prix Is Today: How to Watch the Race Live - CNET - June 20th, 2022 [June 20th, 2022]
- Sentience is the wrong discussion to have on AI right now - TechTalks - June 20th, 2022 [June 20th, 2022]
- Headlines at 10:30 am on 20th June 2022 - The Indian Express - June 20th, 2022 [June 20th, 2022]
- 5 Chess Brilliancies That Stockfish Hates - Chess.com - June 11th, 2022 [June 11th, 2022]
- Carlsen Wins, Leads, Hits A 2870 Live Rating - Chess.com - June 11th, 2022 [June 11th, 2022]
- 21 things to do with kids in San Diego County in June - The San Diego Union-Tribune - June 11th, 2022 [June 11th, 2022]
- Is This Cooling Technology Company Ready To Heat Up? - Benzinga - Benzinga - June 3rd, 2022 [June 3rd, 2022]
- Calendar of events and activities throughout Downriver - Southgate News Herald - June 3rd, 2022 [June 3rd, 2022]
- Tilting Point partners with Polygon on Web3 games - VentureBeat - May 11th, 2022 [May 11th, 2022]
- Online booking agents have been behaving like kings - it's time to topple them - City A.M. - April 17th, 2022 [April 17th, 2022]
- Chess Games - Play Chess Games on CrazyGames - March 29th, 2022 [March 29th, 2022]
- A tale of two universities and two engines - Chess News - March 26th, 2022 [March 26th, 2022]
- Charity Cup: Anton wins three in a row to reach knockout - Chess News - March 26th, 2022 [March 26th, 2022]
- Formula 1: How to Watch the Bahrain Grand Prix and F1 Racing in 2022 - CNET - March 26th, 2022 [March 26th, 2022]
- Praggnanandhaa, 16, becomes only third Indian to beat Magnus Carlsen in stunning upset - ESPN - February 21st, 2022 [February 21st, 2022]
- Is Artificial Intelligence as Intelligent as We Think it is? - Analytics Insight - February 17th, 2022 [February 17th, 2022]
- Didnt Become a Hostage- Former World Chess Champion Calls Magnus Carlsen the Bridge Between Traditional and Modern Chess - EssentiallySports - February 17th, 2022 [February 17th, 2022]
- Can the academy rein in Big Tech? - Times Higher Education - February 17th, 2022 [February 17th, 2022]
- FIDE World Women's Team Championship Final: Russia Wins Gold In Victory Over India - Chess.com - February 17th, 2022 [February 17th, 2022]
- Researchers warn that social media may be fundamentally at odds with science - TechCrunch - February 15th, 2022 [February 15th, 2022]
- Battle of the Sexes: Men triumph! - Chessbase News - February 9th, 2022 [February 9th, 2022]
- Battle of the Sexes: Men increase lead - Chessbase News - February 5th, 2022 [February 5th, 2022]
- Chairman of the board | Boris Starling - The Critic - February 5th, 2022 [February 5th, 2022]
- Using AI in Recruiting - Onrec - February 5th, 2022 [February 5th, 2022]
- Arena Download - Complete GUI for chess engines that will ... - January 24th, 2022 [January 24th, 2022]
- A hundred years of exactitude: Jos Ral Capablanca - TheArticle - January 24th, 2022 [January 24th, 2022]
- Intel Core i5-12400 vs AMD Ryzen 5 5600X Face-Off: The Gaming Value Showdown - Tom's Hardware - January 24th, 2022 [January 24th, 2022]
- software - Why dont chess engines use Node.js? - Chess ... - December 29th, 2021 [December 29th, 2021]
- Stockfish - Chess Engines - Chess.com - December 27th, 2021 [December 27th, 2021]
- Top 10 Strongest Chess Engines In 2021 - Hercules Chess - December 23rd, 2021 [December 23rd, 2021]
- The 10 Greatest Blitz Chess Games Of All Time - Chess.com - December 23rd, 2021 [December 23rd, 2021]
- Ninja, the worlds top streamer, on how video games can make you smarter about money and investing - MarketWatch - December 17th, 2021 [December 17th, 2021]
- 8 Reasons To Play In The 2022 Daily Chess Championship - Chess.com - December 15th, 2021 [December 15th, 2021]
- World Chess Championship - the Arena - Chessbase News - December 7th, 2021 [December 7th, 2021]
- The World Chess Championship Opens With An Endless Knight-Rook Dance - FiveThirtyEight - November 27th, 2021 [November 27th, 2021]
- Play chess: online and computer chess on real boards in the test - Market Research Telecast - November 27th, 2021 [November 27th, 2021]
- The 5 Best Computer Chess Engines - Chess.com - November 15th, 2021 [November 15th, 2021]
- 10 Strongest Free Chess Engines [all above 3000 ELO] at ... - November 15th, 2021 [November 15th, 2021]
- Chessprogramming wiki - November 3rd, 2021 [November 3rd, 2021]
- Stockfish can crush you at chess even more efficiently in the 14.1 update - Neowin - November 3rd, 2021 [November 3rd, 2021]
- Grand Swiss: Shirov and Najer join Firouzja in the lead - Chessbase News - November 3rd, 2021 [November 3rd, 2021]
- Deadmau5's 'Oberhasli' is what it looks like when the metaverse comes for music fans - Mashable South East Asia - October 26th, 2021 [October 26th, 2021]
- Deadmau5's 'Oberhasli' is what it looks like when the metaverse comes for music fans - Mashable - October 24th, 2021 [October 24th, 2021]
- Deep Blue - Chess.com - October 17th, 2021 [October 17th, 2021]
- AlphaZero Crushes Stockfish In New 1,000-Game Match - Chess.com - October 17th, 2021 [October 17th, 2021]
- Free UCI-Compatible Chess Programs for the Stockfish Engine - HobbyLark - October 17th, 2021 [October 17th, 2021]
- CORRECTING and REPLACING RazerCon Is Back for Round II: Tune in for a Keynote By CEO Min-Liang Tan Filled With Exclusive New Announcements and Guest... - September 29th, 2021 [September 29th, 2021]