{"id":1210281,"date":"2022-09-24T14:39:29","date_gmt":"2022-09-24T14:39:29","guid":{"rendered":"https:\/\/www.ghanamma.com\/2022\/09\/24\/multilingual-laughing-pitfall-playing-and-streetwise-ai\/"},"modified":"2022-09-24T14:39:29","modified_gmt":"2022-09-24T14:39:29","slug":"multilingual-laughing-pitfall-playing-and-streetwise-ai","status":"publish","type":"post","link":"https:\/\/www.ghanamma.com\/2022\/09\/24\/multilingual-laughing-pitfall-playing-and-streetwise-ai\/","title":{"rendered":"Multilingual, laughing, Pitfall-playing and streetwise AI \u2022"},"content":{"rendered":"<p><\/p>\n<p id=\"speakable-summary\">Research in the field of machine learning and AI, now a key technology in practically every industry and company, is far too voluminous for anyone to read it all. This column,&nbsp;Perceptron, aims to collect some of the most relevant recent discoveries and papers \u2014 particularly in, but not limited to, artificial intelligence \u2014 and explain why they matter.<\/p>\n<p id=\"speakable-summary\">Over the past few weeks, researchers at Google have demoed an AI system, PaLI, that can perform many tasks in over 100 languages. Elsewhere, a Berlin-based group launched a project called Source+ that\u2019s designed as a way of allowing artists, including visual artists, musicians and writers, to opt into \u2014 and out of \u2014 allowing their work being used as training data for AI.<\/p>\n<p>AI systems like OpenAI\u2019s GPT-3 can generate fairly sensical text, or summarize existing text from the web, ebooks and other sources of information. But they\u2019re historically been limited to a single language, limiting both their usefulness and reach.<\/p>\n<p>Fortunately, in recent months, research into multilingual systems has accelerated \u2014 driven partly by community efforts like Hugging Face\u2019s Bloom. In an attempt to leverage these advances in multilinguality, a Google team created PaLI, which was trained on both images and text to perform tasks like image captioning, object detection and optical character recognition.<\/p>\n<div id=\"attachment_2405492\" class=\"wp-caption aligncenter\">\n<p id=\"caption-attachment-2405492\" class=\"wp-caption-text\"><strong>Image Credits:<\/strong> Google<\/p>\n<\/div>\n<p>Google claims that PaLI can understand 109 languages and the relationships between words in those languages and images, enabling it to \u2014 for example \u2014 caption a picture of a postcard in French. While the work remains firmly in the research phases, the creators say that it illustrates the important interplay between language and images \u2014 and could establish a foundation for a commercial product down the line.<\/p>\n<p>Speech is another aspect of language that AI is constantly improving in. Play.ht recently showed off a new text-to-speech model that puts a remarkable amount of emotion and range into its results. The clips it posted last week sound fantastic, though they are of course cherry-picked.<\/p>\n<p>We generated a clip of our own using the intro to this article, and the results are still solid:<\/p>\n<p><!--[if lt IE 9]&gt;&lt;![endif]--><br \/><audio class=\"wp-audio-shortcode\" id=\"audio-2400656-1\" preload=\"none\" controls=\"controls\">https:\/\/techcrunch.com\/wp-content\/uploads\/2022\/09\/perceptron-peregrine.wav<\/audio><\/p>\n<p>Exactly what this type of voice generation will be most useful for is still unclear. We\u2019re not quite at the stage where they do whole books \u2014 or rather, they can, but it may not be anyone\u2019s first choice yet. But as the quality rises, the applications multiply.<\/p>\n<p>Mat Dryhurst and Holly Herndon \u2014 an academic and musician, respectively \u2014 have partnered with the organization Spawning to launch Source+, a standard they hope will bring attention to the issue of photo-generating AI systems created using artwork from artists who weren\u2019t informed or asked permission. Source+, which doesn\u2019t cost anything, aims to allow artists to disallow their work to be used for AI training purposes if they choose.<\/p>\n<p>Image-generating systems like Stable Diffusion and DALL-E 2 were trained on billions of images scraped from the web to \u201clearn\u201d how to translate text prompts into art. Some of these images came from public art communities like ArtStation and DeviantArt \u2014 not necessarily with artists\u2019 knowledge \u2014 and imbued the systems with the ability to mimic particular creators, including artists like Greg Rutowski.<\/p>\n<div id=\"attachment_2370623\" class=\"wp-caption aligncenter\"><img decoding=\"async\" aria-describedby=\"caption-attachment-2370623\" loading=\"lazy\" class=\"wp-image-2370623 size-full\" src=\"https:\/\/www.ghanamma.com\/gp\/wp-content\/uploads\/2022\/08\/merged-0005.png\" alt=\"Stability AI Stable Diffusion\" width=\"2560\" height=\"512\"><\/p>\n<p id=\"caption-attachment-2370623\" class=\"wp-caption-text\">Samples from Stable Diffusion.<\/p>\n<\/div>\n<p>Because of the systems\u2019 knack for imitating art styles, some creators fear that they could threaten livelihoods. Source+ \u2014 while voluntary \u2014 could be a step toward giving artists greater say in how their art\u2019s used, Dryhurst and Herndon say \u2014 assuming it\u2019s adopted at scale (a big if).<\/p>\n<p>Over at DeepMind, a research team is attempting to solve another longstanding problematic aspect of AI: its tendency to spew toxic and misleading information. Focusing on text, the team developed a chatbot called Sparrow that can answer common questions by searching the web using Google. Other cutting-edge systems like Google\u2019s LaMDA can do the same, but DeepMind claims that Sparrow provides plausible, non-toxic answers to questions more often than its counterparts.<\/p>\n<p>The trick was aligning the system with people\u2019s expectations of it. DeepMind recruited people to use Sparrow and then had them provide feedback to train a model of how useful the answers were, showing participants multiple answers to the same question and asking them which answer they liked the most. The researchers also defined rules for Sparrow such as \u201cdon\u2019t make threatening statements\u201d and \u201cdon\u2019t make hateful or insulting comments,\u201d which they had participants impose on the system by trying to trick it into breaking the rules.<\/p>\n<div id=\"attachment_2405674\" class=\"wp-caption aligncenter\"><img decoding=\"async\" aria-describedby=\"caption-attachment-2405674\" loading=\"lazy\" class=\"size-full wp-image-2405674\" src=\"https:\/\/www.ghanamma.com\/gp\/wp-content\/uploads\/2022\/09\/sparrow-split.jpg\" alt width=\"1024\" height=\"786\"><\/p>\n<p id=\"caption-attachment-2405674\" class=\"wp-caption-text\">Example of DeepMind\u2019s sparrow having a conversation.<\/p>\n<\/div>\n<p>DeepMind acknowledges that Sparrow has room for improvement. But in a study, the team found the chatbot provided a \u201cplausible\u201d answer supported with evidence 78% of the time when asked a factual question and only broke the aforementioned rules 8% of the time. That\u2019s better than DeepMind\u2019s original dialogue system, the researchers note, which broke the rules roughly three times more often when tricked into doing so.<\/p>\n<p>A separate team at DeepMind tackled a very different domain recently: video games that historically have been tough for AI to master quickly. Their system, cheekily called MEME, reportedly achieved \u201chuman-level\u201d performance on 57 different Atari games 200 times faster than the previous best system.<\/p>\n<p>According to DeepMind\u2019s paper detailing MEME, the system can learn to play games by observing roughly 390 million frames \u2014 \u201cframes\u201d referring to the still images that refresh very quickly to give the impression of motion. That might sound like a lot, but the previous state-of-the-art technique required 80 <em>billion <\/em>frames across the same number of Atari games.<\/p>\n<div id=\"attachment_2405524\" class=\"wp-caption aligncenter\"><img decoding=\"async\" aria-describedby=\"caption-attachment-2405524\" loading=\"lazy\" class=\"wp-image-2405524 size-full\" src=\"https:\/\/www.ghanamma.com\/gp\/wp-content\/uploads\/2022\/09\/image-53.webp.png\" alt=\"DeepMind MEME\" width=\"1024\" height=\"338\"><\/p>\n<p id=\"caption-attachment-2405524\" class=\"wp-caption-text\"><strong>Image Credits:<\/strong> DeepMind<\/p>\n<\/div>\n<p>Deftly playing Atari might not sound like a desirable skill. And indeed, some critics argue games are a flawed AI benchmark because of their abstractness and relative simplicity. But research labs like DeepMind believe the approaches could be applied to other, more useful areas in the future, like robots that more efficiently learn to perform tasks by watching videos or self-improving, self-driving cars.<\/p>\n<p>Nvidia had a field day on the 20th announcing dozens of products and services, among them several interesting AI efforts. Self-driving cars are one of the company\u2019s foci, both powering the AI and training it. For the latter, simulators are crucial and it is likewise important that the virtual roads resemble real ones. They describe a new, improved content flow that accelerates bringing data collected by cameras and sensors on real cars into the digital realm.<\/p>\n<div id=\"attachment_2405746\" class=\"wp-caption aligncenter\"><img decoding=\"async\" aria-describedby=\"caption-attachment-2405746\" loading=\"lazy\" class=\"size-full wp-image-2405746\" src=\"https:\/\/www.ghanamma.com\/gp\/wp-content\/uploads\/2022\/09\/drivesim-1.jpg\" alt width=\"1024\" height=\"287\"><\/p>\n<p id=\"caption-attachment-2405746\" class=\"wp-caption-text\">A simulation environment built on real-world data.<\/p>\n<\/div>\n<p>Things like real-world vehicles and irregularities in the road or tree cover can be accurately reproduced, so the self-driving AI doesn\u2019t learn in a sanitized version of the street. And it makes it possible to create larger and more variable simulation settings in general, which aids robustness. (Another image of it is up top.)<\/p>\n<p>Nvidia also introduced its IGX system for autonomous platforms in industrial situations \u2014 human-machine collaboration like you might find on a factory floor. There\u2019s no shortage of these, of course, but as the complexity of tasks and operating environments increases, the old methods don\u2019t cut it any more and companies looking to improve their automation are looking at future-proofing.<\/p>\n<div id=\"attachment_2406640\" class=\"wp-caption aligncenter\"><img decoding=\"async\" aria-describedby=\"caption-attachment-2406640\" loading=\"lazy\" class=\"size-full wp-image-2406640\" src=\"https:\/\/www.ghanamma.com\/gp\/wp-content\/uploads\/2022\/09\/edge-ai-promo-igx-launch-press-1260x680-1.jpg\" alt width=\"1024\" height=\"553\"><\/p>\n<p id=\"caption-attachment-2406640\" class=\"wp-caption-text\">Example of computer vision classifying objects and people on a factory floor.<\/p>\n<\/div>\n<p>\u201cProactive\u201d and \u201cpredictive\u201d safety are what IGX is intended to help with, which is to say catching safety issues before they cause outages or injuries. A bot may have its own emergency stop mechanism, but if a camera monitoring the area could tell it to divert before a forklift gets in its way, everything goes a little more smoothly. Exactly what company or software accomplishes this (and on what hardware, and how it all gets paid for) is still a work in progress, with the likes of Nvidia and startups like Veo Robotics feeling their way through.<\/p>\n<p>Another interesting step forward was taken in Nvidia\u2019s home turf of gaming. The company\u2019s latest and greatest GPUs are built not just to push triangles and shaders, but to quickly accomplish AI-powered tasks like its own DLSS tech for uprezzing and adding frames.<\/p>\n<p>The issue they\u2019re trying to solve is that gaming engines are so demanding that generating more than 120 frames per second (to keep up with the latest monitors) while maintaining visual fidelity is a Herculean task even powerful GPUs can barely do. But DLSS is sort of like an intelligent frame blender that can increase the resolution of the source frame without aliasing or artifacts, so the game doesn\u2019t have to push quite so many pixels.<\/p>\n<p>In DLSS 3, Nvidia claims it can generate entire additional frames at a 1:1 ratio, so you could be rendering 60 frames naturally and the other 60 via AI. I can think of several reasons that might make things weird in a high performance gaming environment, but Nvidia is probably well aware of those. At any rate you\u2019ll need to pay about a grand for the privilege of using the new system, since it will only run on RTX 40 series cards. But if graphical fidelity is your top priority, have at it.<\/p>\n<div id=\"attachment_2406708\" class=\"wp-caption aligncenter\"><img decoding=\"async\" aria-describedby=\"caption-attachment-2406708\" loading=\"lazy\" class=\"size-full wp-image-2406708\" src=\"https:\/\/www.ghanamma.com\/gp\/wp-content\/uploads\/2022\/09\/remote-forest_1663772841389_x2.jpg\" alt width=\"1024\" height=\"619\"><\/p>\n<p id=\"caption-attachment-2406708\" class=\"wp-caption-text\">Illustration of drones building in a remote area.<\/p>\n<\/div>\n<p>Last thing today is a drone-based 3D printing technique from Imperial College London that could be used for autonomous building processes sometime in the deep future. For now it\u2019s definitely not practical for creating anything bigger than a trash can, but it\u2019s still early days. Eventually they hope to make it more like the above, and it does look cool, but watch the video below to get your expectations straight.<\/p>\n<p>[embedded content]<\/p>\n","protected":false},"excerpt":{"rendered":"<p>Research in the field of machine learning and AI, now a key technology in practically every industry and company, is far too voluminous for anyone to read it all. This column,&nbsp;Perceptron, aims to collect some of the most relevant recent discoveries and papers \u2014 particularly in, but not limited to, artificial intelligence \u2014 and explain [&hellip;]<\/p>\n","protected":false},"author":1,"featured_media":1210282,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[21],"tags":[],"class_list":["post-1210281","post","type-post","status-publish","format-standard","has-post-thumbnail","category-celebrity-gossip"],"_links":{"self":[{"href":"https:\/\/www.ghanamma.com\/2022\/wp-json\/wp\/v2\/posts\/1210281","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.ghanamma.com\/2022\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.ghanamma.com\/2022\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.ghanamma.com\/2022\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/www.ghanamma.com\/2022\/wp-json\/wp\/v2\/comments?post=1210281"}],"version-history":[{"count":0,"href":"https:\/\/www.ghanamma.com\/2022\/wp-json\/wp\/v2\/posts\/1210281\/revisions"}],"wp:attachment":[{"href":"https:\/\/www.ghanamma.com\/2022\/wp-json\/wp\/v2\/media?parent=1210281"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.ghanamma.com\/2022\/wp-json\/wp\/v2\/categories?post=1210281"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.ghanamma.com\/2022\/wp-json\/wp\/v2\/tags?post=1210281"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}