<?xml version="1.0" encoding="utf-8"?>
<!-- generator="Joomla! - Open Source Content Management" -->
<?xml-stylesheet href="/plugins/system/jce/css/content.css?badb4208be409b1335b815dde676300e" type="text/css"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
	<channel>
		<title>blog - MUHAI</title>
		<description><![CDATA[Meaning and Understanding
in Human-centric AI]]></description>
		<link>https://muhai.org/blog/11-understanding-everyday-activities</link>
		<lastBuildDate>Tue, 14 Oct 2025 13:40:54 +0000</lastBuildDate>
		<generator>Joomla! - Open Source Content Management</generator>
		<atom:link rel="self" type="application/rss+xml" href="https://muhai.org/blog/11-understanding-everyday-activities?format=feed&amp;type=rss"/>
		<language>en-gb</language>
		<item>
			<title>Can Robots Cook? Culinary challenges for advancing artificial intelligence</title>
			<link>https://muhai.org/blog/11-understanding-everyday-activities/271-can-robots-cook</link>
			<guid isPermaLink="true">https://muhai.org/blog/11-understanding-everyday-activities/271-can-robots-cook</guid>
			<description><![CDATA[<p><img src="https://muhai.org/images/article/blog-cooks-3.png" /></p><h3>&nbsp;</h3>
<h3>Alexane Jouglar, University of Namur.</h3>
<h3><img src="https://muhai.org/images/article/blog-cooks-1.png" alt="blog cooks 1" width="420" height="238" style="display: block; margin-left: auto; margin-right: auto;" />&nbsp;&nbsp;</h3>
<p style="text-align: center;"><span style="font-size: 10pt;">Source: <a href="https://dictionary.cambridge.org/dictionary/english/cook">https://dictionary.cambridge.org/dictionary/english/cook</a></span></p>
<p>Cooking is an act that we perform in our everyday lives to produce (delicious) dishes ready to eat. Many machines exist to help humans cook: slow cookers, multi-cookers, pressure cookers… All these machines are really helpful, but they need the direct intervention of a human. Currently, It is not possible to give a recipe to a machine, place the machine in the kitchen, and ask it to cook the dish. The reason for this is that recipes include a lot of background and implicit knowledge that humans gradually acquire by practicing and interacting with the world and that cannot be understood from the only lecture of the recipe. Let’s take an example.</p>
<p><img src="https://muhai.org/images/article/blog-cooks-2.png" alt="blog cooks 2" width="746" height="786" style="display: block; margin-left: auto; margin-right: auto;" /></p>
<p style="text-align: center;"><span style="font-size: 10pt;">Source: <a href="https://chocolatecoveredkatie.com/vegan-chocolate-chip-cookies-recipe/">https://chocolatecoveredkatie.com/vegan-chocolate-chip-cookies-recipe/</a></span></p>
<p>As we can see in the figure above, the recipe instructions start with “Combine all dry ingredients in a bowl”. As humans, we know what dry ingredients are. We also know that by “all”, the writer means “from all the ingredients that are listed above”. For a machine, this is a lot harder. It needs to be equipped with a lot of background knowledge. That's precisely the challenge posed by a groundbreaking new benchmark introduced in a recent paper published at LREC-COLING 2024:</p>
<p><em>Nevens, J., De Haes, R., Ringe, R., Pomarlan, M., Porzel, R., Beuls, K., &amp; Van Eecke, P.&nbsp;(Accepted/In press).&nbsp;A Benchmark for Recipe Understanding in Artificial Agents. In&nbsp;The 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation</em></p>
<p>This benchmark is the result of a collaboration between the Vrije Universiteit Brussels (VUB), and the University of Bremen (UBremen) of the University of Namur (UNamur). It aims to evaluate whether artificial agents can understand and perform everyday cooking activities. It includes several components:</p>
<ul>
<li>Corpus of Recipes: A collection of 30 recipes of varying complexity.</li>
<li>Procedural Semantic Representation Language: This language helps formalize 38 cooking actions in a way that machines can understand. It breaks down recipes into precise steps and provides a framework for representing cooking processes.</li>
<li>Kitchen Simulators: Both qualitative and quantitative simulators recreate kitchen environments where agents can practice cooking. These simulators simulate the physical aspects of cooking, such as ingredient interactions and cooking times.</li>
<li>Evaluation Procedure: A standardized method for evaluating agent performance.</li>
</ul>
<p>The main task of the benchmark is to translate natural language recipes into a series of cooking actions that can be executed in the simulated kitchen to produce the desired dish. This translation process requires the agent to reason over multiple factors, including the recipe text, the state of the simulated kitchen, common-sense knowledge, and domain-specific cooking knowledge.</p>
<p>Success in this benchmark requires a combination of natural language processing and situated reasoning. By mastering the art of cooking, artificial agents can become more versatile and helpful in various real-world scenarios.</p>
<p><img src="https://muhai.org/images/article/blog-cooks-3.png" alt="blog cooks 3" width="704" height="396" style="display: block; margin-left: auto; margin-right: auto;" /></p>
<p style="text-align: center;"><span style="font-size: 10pt;">Source: Canva</span></p>
<p>The introduction of this benchmark represents a significant step forward in the field of artificial intelligence. It challenges researchers to develop systems that can understand natural language and apply that understanding to practical tasks like cooking. As these systems improve, they have the potential to revolutionize various aspects of our lives, from personal assistance in the kitchen to industrial food production.</p>
<p>Discover the full paper in our dedicated <a href="https://muhai.org/papers">section</a>:&nbsp;Jens Nevens, Robin de Haes, Rachel Ringe, Mihai Pomarlan, Robert Porzel, Katrien Beuls and Paul van Eecke. A Benchmark for Recipe Understanding in Artificial Agents. In Nicoletta Calzolari, Min-Yen Kan, Veronique Hoste, Alessandro Lenci, Sakriani Sakti and Nianwen Xue (eds.).&nbsp;<i>Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024)</i>. 2024, 22–42</p>]]></description>
			<category>Understanding Everyday Activities</category>
			<pubDate>Thu, 09 May 2024 09:09:59 +0000</pubDate>
		</item>
		<item>
			<title>Anaphora Unveiled: Tracking Culinary Transformation in the Tech-Driven Kitchen</title>
			<link>https://muhai.org/blog/11-understanding-everyday-activities/255-anaphora-unv</link>
			<guid isPermaLink="true">https://muhai.org/blog/11-understanding-everyday-activities/255-anaphora-unv</guid>
			<description><![CDATA[<p><img src="https://muhai.org/images/article/Blog3.jpg" /></p><h3>&nbsp;</h3>
<h3>Anna Morbiato, Venice International University.</h3>
<h3><img src="https://muhai.org/images/article/Blog3.jpg" alt="Blog3" width="749" height="499" style="display: block; margin-left: auto; margin-right: auto;" />&nbsp;&nbsp;</h3>
<p>In the dazzling realm where cutting-edge technology meets the world of culinary arts, precision is not just a preference – it's the secret sauce to success. Imagine a futuristic kitchen where robotic chefs whip up gourmet delights with a mere tap on a touchscreen. Now, in this technologically driven culinary landscape, the unsung hero emerges: accurate tracking of coreference. Like a culinary GPS, it navigates all complex recipe transformations undergone by ingredients, ensuring that every ‘it’ and ‘they’ (as in ‘put it in the fridge’ or ‘they are ready when a fork slides in easily’) points to the right flavour-packed entity. Let's explore why this seemingly subtle linguistic skill is the backbone of a seamless, tech-infused gastronomic experience.<br />Take a simple Italian pasta recipe step: 'Boil a pot of water, add some salt, then add your favourite pasta type. Once it's cooked, drain it and then add it directly to the tomato sauce.' While us readers have no doubt what 'it' refers to, it is not as straightforward for a machine to understand that 'it' refers to the pasta and not to other nouns like 'pot,' 'water,' or 'salt.' Still, it is crucial that the right item is added to the sauce for a successful execution of the recipe.</p>
<p><br />Accurate tracking of coreference becomes crucial in a technologically driven culinary landscape. But this comprehension comes with its set of challenges. Recipes are full of pronouns and other ways to refer to items, ingredients, and tools. However, these are sometimes far from transparent and often include not only pronouns, but also even zero anaphors (Ø) (namely using no words at all!). The same sentence as above could work also with no 'it': Boil a pot of water, put some salt, then add your favourite pasta type. Once Ø cooked, drain Ø and then add Ø directly to the tomato sauce.’ Not only ingredients, but also intermediate products of each step (resultant objects, like a mixture or a cream) are frequently left unmentioned, resulting in potential ambiguity. This phenomenon is not marginal; in fact, zero anaphors (also referred to as zero pronouns, null elements, or implicit entities) are more frequent in recipes as compared to other genres. This is especially evident in pro-drop languages like Italian and topic-drop languages like Chinese, where zero anaphora is widespread and often requires inference for correct interpretation. Further challenges include partial coreference phenomena (also known as bridging) as well as evolutive anaphors, namely linguistic expressions that refer to an entity while also indicating a change in its properties or attributes over time. Evolutive anaphors are very common in culinary narratives, where they typically refer to an ingredient or component of the recipe which, in the meantime, has undergone some transformation. <br />Imagine we want to make baked potatoes. The recipe starts by telling us ‘Wask and peel the potatoes.’ Then it goes on with the cooking procedure, which includes garlic, warm milk and butter. The last passage says: “Blend with a potato masher or electric mixer until potatoes are smooth and creamy.” For sure, the entity denoted by the term potatoes in the ingredients and that denoted by the same word in the last passage are very different. Then why do we still use the word potatoes? And why does it sound odd if we do the same in a sentence like ‘Juice the apples, then put them onto a pan.’?</p>
<p><br />Unravelling the complexities of anaphoric references in recipes is the objective of the paper ‘Pragmatics as the secret ingredient for NLP: a cross-linguistic study of reference tracking and evolutive anaphors’. The paper provides an in-depth exploration of coreference encoding and resolution in the domains of linguistics and Natural Language Processing, with a particular emphasis on zero pronouns, partial coreference phenomena, and evolutive anaphors. Additionally, it presents an analysis of the flow of information in recipe texts, as well as of the type, frequency, and nature of anaphoric devices used in authentic recipes drawn from food blogs. It considers texts in English, Italian, and Chinese, which significantly vary in terms of coreference tracking mechanisms and demonstrate different levels of dependence on zero anaphors and inferential processes. The paper gives particular attention to evolutive anaphors and the role played by inference and world knowledge in coreference disambiguation. Recall the pasta recipes above: the pronoun it involves a change in properties of the pasta, from being hard to being soft after cooking. Crucially, inference plays a significant role in coreference disambiguation, in addition to, or in place of, lexical and grammatical encoding. In fact, it is world knowledge and common sense that enable the reader to understand that the anaphors refer to pasta, and not to pot, water, or salt. Any AI systems aiming to effectively address anaphora resolution and achieve complete textual comprehension must integrate inferential processes at a certain stage.</p>]]></description>
			<category>Understanding Everyday Activities</category>
			<pubDate>Thu, 14 Mar 2024 17:05:49 +0000</pubDate>
		</item>
		<item>
			<title>From Kitchen to AI: A Task-based Metric for Measuring Trust </title>
			<link>https://muhai.org/blog/11-understanding-everyday-activities/217-a-task-based-metric-for-measuring-trust-i-from-kitchen-to-ai</link>
			<guid isPermaLink="true">https://muhai.org/blog/11-understanding-everyday-activities/217-a-task-based-metric-for-measuring-trust-i-from-kitchen-to-ai</guid>
			<description><![CDATA[<p><img src="https://muhai.org/images/pexels-kindel-media-9028917.jpg" /></p><h3>&nbsp;</h3>
<h3>Robert Porzel, University Bremen.</h3>
<p><strong><img src="https://muhai.org/images/BlogRobert_Trust.png" alt="" /></strong></p>
<h3>&nbsp;</h3>
<p><strong>Trust is an important factor in</strong> <strong>human-centric artificial intelligence</strong> – especially for the success and effectiveness of a collaborative task in which the participants rely on each other to achieve specific sub-goals. For example, in household environments, such as a kitchen, mistakes can be made by either party that could not only lead to failure to complete the task, but even to injury through various hot or sharp appliances. Trust in a new system or technology is critical to its success, since people tend to employ systems that they trust, and reject systems that they do not trust.</p>
<p>In the last few years, artificial agents, such as vacuuming robots, have become more common in household environments, and assistants for more complex tasks as cooking or cleaning are being developed. To ensure that these new systems will be accepted, it is important to explore how much people trust an autonomous system to handle these tasks, how this trust changes during use and what factors lead to an increase or decrease in trust. Toward that goal, it is important to find applicable measures for trust. For this, we propose measuring a user's trust in an artificial collaborator during cooperative cooking tasks by analysing the tasks delegated to the artificial partner during collaborative execution of a recipe.</p>
<p>As the delegation of tasks among humans relies on trust, we propose that the tasks given to the artificial collaborator, e.g. while preparing a meal, can supply information on the level of trust the human has in them. If the human assigns intricate, dangerous or important tasks to the artificial agent, e.g. heating or cutting an ingredient, this would indicate, that they trust this partner to complete the task successfully. Should they only delegate minor tasks to the robot – for example, wiping the counter – it indicates, that the robotic partner is only trusted to fulfil simple tasks where errors could easily mitigated. Toward the goal of measuring trust based on task delegation, three different aspects of a task that could influence a human’s tendency to delegate it were chosen in our approach: difficulty, risk and possibility for error mitigation. In addition, it was deemed relevant if a human would supervise the artificial collaborator during a task or even intervene. In addition discount factors were considered that might convince a human to assign a task to a robot even though they do not completely trust the robot, e.g. tediousness of a task or inability to complete a task themselves. These aspects were combined into a basis for a scale, that can be used to determine the level of trust the human put into the artificial partner when delegating this specific task to them.</p>
<p>To observe humans during cooperative cooking with an artificial partner, a VR application was developed in the Unity game engine for use with an Oculus Quest HMD. In this application the user is placed in a kitchen environment together with a virtual robot. The user can interact with various objects in the kitchen by grabbing them with either their hands or the controllers and then complete various cooking tasks by moving them in appropriate ways -- e.g., moving a whisk in circular motions through a bowl containing the different ingredients to be mixed. In addition, the user can order the robot to fulfil any of the needed cooking tasks for recipe completion -- e.g., portioning a certain&nbsp;amount of an ingredient into a bowl -- or some supporting such as cleaning, tidying or fetching objects for the user. For these orders a delegation-type interface is used, where the user orders the robot to fulfil a task in a declarative manner, but is not required to give details on how the task should be completed.</p>
<p>In the future this metric could be part of a bigger set of measures for trust specific to cooperative tasks, that includes other aspects such as the phrasing of orders given to the robot. Similarly, it could be modified for further household tasks, that could in the future be assigned to household robots. Predictions made by a graphical model based on these metrics could also be used to adjust robot behavior at runtime to calibrate trust to the appropriate level for optimal cooperation. The described test environment and scale could be used in the future to explore different robot appearances and behaviors and how they affect trust, as well as trust development over time when the human can observe the robot complete tasks successfully or make errors.</p>]]></description>
			<category>Understanding Everyday Activities</category>
			<pubDate>Tue, 23 May 2023 09:44:37 +0000</pubDate>
		</item>
		<item>
			<title>Narrative Objects</title>
			<link>https://muhai.org/blog/11-understanding-everyday-activities/207-narrative-objects</link>
			<guid isPermaLink="true">https://muhai.org/blog/11-understanding-everyday-activities/207-narrative-objects</guid>
			<description><![CDATA[<p><img src="https://muhai.org/images/article/Narrative_objects_MUHAI.jpeg" /></p><h3><a href="https://muhai.org/people"><br />Mihai Pomarlan, UHB.</a></h3>
<p><img src="https://muhai.org/images/article/Narrative_objects_MUHAI.jpeg" alt="Narrative objects MUHAI" width="1600" height="1312" /></p>
<p>A recurring theme of AI research has turned out to be that what should be easy often is not. Consider this question&nbsp;– "what can I cut a stick of butter with?" You probably already thought of an answer, and if pressed, you could invent more creative ones. A string might do, or the edge of a glass perhaps if no knife is available&nbsp;– though you might protest if I suggested a jar's edge. It will be annoying to get the butter rests out of the threading.<br /> <br /> You can answer such questions as the one above, and even employ some creativity in the answers, because, presumably, you have built a sophisticated model of the world and the things in it through your experience of interaction with them. You have an idea of what an object can do, and how it will change in various situations. Now that you have this model, it is second nature to reason with it. Could we implement something similar inside a robot to help us with housework? As always, let's start with moderate ambition. Could we implement a model of "normal" object use, so that our housework robot could at least give the (to us) obvious answer? We'll worry about whether a robot should be creative, and if so how, some other time.<br /> <br /> But if all we want is a database of normal object uses, surely we have such a thing already. You could google "what can I cut a stick of butter with?" and get good answers. If google can answer it, surely a robot also can. And yet ... "what should be easy often is not". It's up to you to select which of the answers are relevant to you, so finding a result with google is often more of a collaboration than it first appears&nbsp;– and both partners should have some idea of the thing searched for. When I ran this experiment, the first article google found was about "cutting in butter", a technique to mix in butter with the dry ingredients for baking. This is not quite what I meant, so I can go down the list of search results to find one that tells me what I expected to find. Would a robot who did not already know how butter and cutting work be able to use these results? And what google finds is in natural language, rather than a nice, well-structured, immediately usable message for a computer program.<br /> <br /> So let's say google is a bit too tricky for our robot, but surely there are large databases out there which cover this sort of everyday commonsense knowledge. Indeed, since commonsense reasoning is recognized as a challenge in AI, there are many efforts towards building commonsense knowledge resources&nbsp;– and this time, the data is represented in machine readable format. Querying and interpreting the results would be easy for a robot to do. The problem is that, were the robot to ask our cutting butter question, he's likely to not get any answer. It turns out, existing commonsense knowledge resources are mostly useful to assisst search engines towards organizing the websites they crawl through into a more semantically rich web. For example, a commonsense knowledge resource might know of many famous people, and that they were human beings, and that human beings have birth dates. Or it would know of countries, and their capital cities. Such information is useful to construct a brief summary of knowledge about an entity, and will also assist in retrieving more documents about that entity&nbsp;– but now we are again in the area of a search engine providing results for a human being to select from. And the human being asked for the search. "What should be easy often is not"&nbsp;– in this case, because if it is easy for a human there is no point in firing up a search engine for an answer. As a result, what should be a commonsense knowledge resource ends up being a resource for trivial questions instead.<br /> <br /> This does not make existing commonsense knowledge resources useless to our robot, but we need to add some more knowledge to them first. Some of this object usability knowledge was obtained by colleagues working on a related project, which involved the development of games through which human players could reveal their preferred object combinations when performing various tasks. Some valuable information we obtained from linguistic resources that describe the roles objects can play in various kinds of events, and restrictions on what objects can play which roles. Finally, we used a cognitively motivated ontological foundation to construct our model upon, which now allows the robot to reason with it and answer questions such as "what can this object be used for?", "can this object be used for a particular task?", and "what can this object be used with when performing a task?''. Altogether, these are the existing knowledge resources we have used so far are: the CommonSense Knowledge Graph (CSKG), the linguistic resources VerbNET, WordNET, the data from the game with a purpose ToolFeud that our colleagues collected, the DOLCE Ultra Lite foundational ontology, and the SOcio-physical Model of Activity (SOMA) that we are developping in a related project. After also some manual input and corrections of data collected semi-automatically from CSKG, we obtained a new knowledge resource, SOMA_DFL.<br /> <br /> There is always more work to be done, of course. Part of that work, ongoing at the moment, is to incorporate some causal knowledge in the SOMA_DFL knowledge base: what happens to an object if some action is performed upon it, how does the quality of an object affect the event it participates in? None of this refers to any advanced knowledge of science; what we want to have in SOMA_DFL is the kind of knowledge that is so useful, and so obvious&nbsp;– to human beings&nbsp;– that almost no one thinks worth writing down. What should be easy often is not&nbsp;– a robot trying to do houeswork will need that knowledge, even as it lacks a natural intuition for it.<br /> <br /> In any case, SOMA_DFL is available at the repository&nbsp;<span style="text-decoration: underline;"><a href="https://github.com/ease-crc/ease_lexical_resources" target="blank">https://github.com/ease-crc/ease_lexical_resources</a></span>. If your robots want to know how to choose tools for some job, give it a try. If they don't find an answer, let us know!</p>
<p><strong>Credits<br /></strong>Intro image - Photo by Stoica Ionela via <span style="text-decoration: underline;"><a href="https://unsplash.com/photos/h26wHZ03fjA" target="blank">Unsplash</a></span><strong><br /></strong></p>]]></description>
			<category>Understanding Everyday Activities</category>
			<pubDate>Thu, 24 Nov 2022 10:56:36 +0000</pubDate>
		</item>
		<item>
			<title>Deep Understanding of Everyday Activity Commands</title>
			<link>https://muhai.org/blog/11-understanding-everyday-activities/198-deep-understanding-of-everyday-activity-commands</link>
			<guid isPermaLink="true">https://muhai.org/blog/11-understanding-everyday-activities/198-deep-understanding-of-everyday-activity-commands</guid>
			<description><![CDATA[<p><img src="https://muhai.org/images/article/UHB_Deep_Understanding_of_Everyday_Activity_Commands.jpg" /></p><h3><br />Robert Porzel, Rainer Malaka, et al., UHB.&nbsp;</h3>
<p><img src="https://muhai.org/images/article/Blog_Deep_Understanding_1560x1280.jpg" alt="UHB Deep Understanding of Everyday Activity Commands" width="1560" height="1280" /></p>
<p>Performing household activities such as cooking and cleaning have, until recently, been the exclusive provenance of human participants. However, the development of robotic agents that can perform different tasks of increasing complexity is slowly changing this state of affairs, creating new opportunities in the domain of household robotics. Most commonly, any robot activity starts with the robot receiving directions or commands for that specific activity. Today, this is mostly done via programming languages or pre-defined user interfaces, but this changes rapidly.</p>
<p>For everyday activities, instructions could be given either verbally from a human or through written texts such as recipes and procedures found in online repositories, e.g., from wikiHow. From the perspective of the robot, these instructions tend to be vague and imprecise as natural language generally employs ambiguous, abstract, and non-verbal cues. Often, instructions omit vital semantic components such as determiners, quantities, or even the objects they refer to.</p>
<p>Still, asking a human to, for instance, "take the cup to the table" will typically result in a satisfactory outcome. Humans excel despite many uncertain variables existing in the environment: for example, the cups might all have been put away in a cupboard or the path to the table could be blocked by chairs.</p>
<p>In comparison, artificial agents lack the same depth of symbol grounding between linguistic cues and real world objects, as well as the capacity for insight and prospection to reason about instructions in relation to the real world and the changing states of that world. In order to turn an underspecified text, issued or taken from the web, into a detailed robotic action plan, various processing steps are necessary, some based on symbolic reasoning and some on numeric simulations or data sets.</p>
<p>The research question the MUHAI project, therefore, encompasses the following question: how can ontological knowledge be used to extract and evaluate parameters from a natural language direction in order to simulate it formally? The solution proposed in this work is the Deep Language Understanding (DLU) approach. This approach employs the ontological Socio-physical Model of Activities (SOMA), which serves not only to define the interfaces of the multi-component pipeline but also to connect numeric data and simulations with symbolic reasoning processes.</p>
<p>Directions, instructions delivered textually, and commands, delivered verbally, are special instructions as they always demand an action. In contrast, other instructions can also be descriptions or specifications of circumstances. For example, "take the cup to the table" demands the addressee to perform an action. Sentences such as "knives must be placed to the right of the plate" or "the onions should be brown after 20 minutes" might entail directions or commands, but they do not explicitly call to action.</p>
<p>Here, the focus lies on directions as an important special case of instructions. More specifically, we first look at directions in which a trajector needs to be moved by an agent along a trajectory, thus directions which can be described using a Source-Path-Goal schema. For this we integrated a natural language understanding system which is capable of simulating natural language directions using a formal ontology as a common interface for individual components. Through this work we showed that the architecture is successful in simulating an everyday activity task.</p>
<p><strong>Illustration:<br /><img src="https://muhai.org/images/article/Deep_Understanding_of_Everyday_Activity_Commands.png" alt="Deep Understanding of Everyday Activity Commands" width="1579" height="1405" /><br /></strong></p>
<p><strong>Online Demonstration:<br /></strong>DLU runs on a live instance at <span style="text-decoration: underline;"><a href="https://litmus.informatik.uni-bremen.de/dlu/">https://litmus.informatik.uni-bremen.de/dlu/</a></span></p>
<p><strong>Video Demonstration:<br /></strong>The video is available at <span style="text-decoration: underline;"><a href="https://osf.io/t2mnw/">https://osf.io/t2mnw/</a></span></p>
<p>
<video src="https://muhai.org/images/video/DLU-video.mp4" controls="controls" width="1920" height="1080"></video>
</p>
<p>&nbsp;</p>
<p><strong>Open Data:<br /></strong>The source code, software, and hardware information used to build and run DLU are available at <span style="text-decoration: underline;"><a href="https://osf.io/nbxsp/">https://osf.io/nbxsp/</a></span></p>
<p>The ontologies and their metrics are available at <span style="text-decoration: underline;"><a href="https://osf.io/e7uck/">https://osf.io/e7uck/</a></span><br /><br /><strong>Credits&nbsp;<br /></strong>Intro photo created by&nbsp;<span style="text-decoration: underline;"><a href="https://www.freepik.com/photos/kitchen-cupboard">kjpargeter - www.freepik.com</a></span></p>]]></description>
			<category>Understanding Everyday Activities</category>
			<pubDate>Thu, 24 Mar 2022 10:54:54 +0000</pubDate>
		</item>
		<item>
			<title>Curiosity-Driven Exploration of Pouring Liquids</title>
			<link>https://muhai.org/blog/11-understanding-everyday-activities/193-curiosity-driven-exploration-of-pouring-liquids</link>
			<guid isPermaLink="true">https://muhai.org/blog/11-understanding-everyday-activities/193-curiosity-driven-exploration-of-pouring-liquids</guid>
			<description><![CDATA[<p><img src="https://muhai.org/images/article/curiosity_pouring_liquids.jpg" /></p><h3><a href="https://muhai.org/aboutus/people"><br />Mihai Pomarlan, UHB.<br /><br /></a><img src="https://muhai.org/images/article/curiosity_pouring_liquids.jpg" alt="curiosity pouring liquids" width="780" height="640" /></h3>
<p>Babies, puppies, kittens may be bundles of joy but they are also agents of pure chaos. They knock things over, stick their fingers where they're not supposed to, and get a taste or sniff of anything they can, all just for fun of course. Or is it "just for fun"? In the moment, it is, but this playfulness of infants also allows them to build the intuitive models of the physical world which are needed to cope with that world "seriously".</p>
<p>The paragraph above is not news to anyone who has witnessed the development of a young mind, but it does highlight several questions that are still cutting-edge in both cognitive science and robotics. What is an "intuitive" model of the world? How would such a thing be acquired? Roboticists add a further question -- how could such models be put to use to help robots cope with the open-ended, loosely structured world of human existence?</p>
<p>In the work we summarize here we have only begun to scratch at such issues, and so a tempering of expectations is honest, necessary -- but also illustrative of both the problems we face and the methods we choose to attack them with. Let us start with the term "curiosity", or greed for the new as the Germans would say. What does it mean to be curious?</p>
<p>Curiosity -- as greed for new, not as strangeness -- is something that babies or kittens can have, but not rocks. That is, curiosity is a quality that an "agent" can have, that is, some being interacting in a purposeful way with its environment. In science and engineering, we interpret "purposeful way of interacting" to mean "acting in such a way so as to optimize a goal function". We can also measure information; we can measure information <em>gain</em>. So, in scientist/engineering terms, "curiosity" is going to mean some version of, "acting in such a way so as to maximize information gain about one's environment".</p>
<p>With some charity, this even seems like a workable definition of human curiosity in general, but it does obscure a complication: no agent has time to make sense of all the information the environment throws at it. Therefore, making sense of the world at any particular moment involves choosing what to look at and discarding the rest, and this human (and presumably, feline, canine etc.) ability to readjust interests remains, as far as cognitive science goes, magic. We do not have a good theory for it, we do not have a way to implement it computationally. We did not try to do it in our work either.</p>
<p>We therefore set our sights on a more modest goal than the open-ended human curiosity. Call it a kind of closed, "robot" curiosity if you will. Our simulated agent is given a well-defined task -- pouring a liquid from one bowl to another -- and a fixed set of "output parameters" to judge the task by -- such as, how much of the liquid reaches the destination, does the destination container move etc. The purpose of the agent is to learn how variations in manner of pouring affect the output parameters. Note, the simulated agent's purpose is not to pour well! Indeed, we want it to pour badly sometimes because mistakes can be informative of how the physical act of pouring works. Besides, no one will cry over simulated spilt milk.</p>
<p>And so our agent pours the liquid, from one bowl to another, in slightly different ways each time, and building a "model" of pouring as it does so. This "model" is like a collection of rules of thumb for predicting outcomes, such as "pouring from too high will produce spillage". We also allow the agent to control physical parameters of the bowls and liquid as well -- what happens if the liquid is very viscous, or very bouncy? All together, there are several billion ways of pouring that our agent could simulate, which is several billion simulations too many to be practical. This is where the "robot curiosity" kicks in.</p>
<p>Rather than try out all variations in pouring, or trying them randomly, the agent uses the pouring model it constructed to choose what to try next. Specifically, it looks at the model to identify where it is not too confident in making predictions -- varying the manner of pouring in those directions is likely to produce behaviors not seen before! Hence, our agent acts -- adjusts the manner of pouring and the physical parameters of the involved objects -- in ways it expects to result in more information -- new behaviors, new rules of thumb to add to its model.</p>
<p>With only a few hundred simulations, rather than several billion, it can arrive at such sensible prediction rules as "pouring from too high/far to the side will produce spillage" and "a viscous fluid needs time to pour out of the source container". Nothing world-shattering -- this sort of stuff is obvious to us humans. Then again, we have the benefit of a truly curious childhood and a better brain. For an agent that was not given any knowledge of pouring and had to start from scratch -- aside from the definition of possible variations and output parameters to look at of course --, ours is not doing too badly.</p>
<p>Much work remains to be done however. The accuracy of the agent's model depends a lot on the methods it uses to group together sets of numerical parameters under the same qualitative umbrella; loosely speaking, this is what we do when we would say someone is tall if their height is sufficiently large. Currently, this conversion between numerical sets and qualitative values is done by fixed mappings, and it should instead be learned as well. But this is a topic of a future work and, we hope, a future post here!</p>
<p><strong>Credits</strong><br />Photo by <a href="https://muhai.org/utm_content=attributionCopyText&amp;utm_medium=referral&amp;;utm_source=pexels">MoldyVintage Photo</a> via&nbsp;<strong><a href="https://www.pexels.com/it-it/foto/caffe-impilati-tazze-scrosciante-8932524/?utm_content=attributionCopyText&amp;utm_medium=referral&amp;utm_source=pexels">Pexels</a></strong></p>]]></description>
			<category>Understanding Everyday Activities</category>
			<pubDate>Tue, 15 Feb 2022 10:22:23 +0000</pubDate>
		</item>
		<item>
			<title>Toward a formal theory of narratives</title>
			<link>https://muhai.org/blog/11-understanding-everyday-activities/188-toward-a-formal-theory-of-narratives</link>
			<guid isPermaLink="true">https://muhai.org/blog/11-understanding-everyday-activities/188-toward-a-formal-theory-of-narratives</guid>
			<description><![CDATA[<p><img src="https://muhai.org/images/article/unnamed.jpg" /></p><h3><a href="https://muhai.org/aboutus/people"><br />Robert Porzel, UHB.</a></h3>
<p><img src="https://muhai.org/images/article/unnamed.jpg" alt="nik shuliahin C0re10PHPfQ unsplash" width="1560" height="1280" /></p>
<p>The activities of people as well as of artificial agents in reality, virtual reality or simulation can be recorded as data that discretize trajectories of body parts and the ensuing force events. While these data provide vast amounts of information they are, by themselves, meaningless. Only when we put them into context we assign a specific meaning to these data. Increasingly, the notion of narratives is being used to describe the result of this semiotic process, i.e. we observe events and fit them into a story that makes sense to us.</p>
<p>The concept of a narrative has migrated from its original domain in the literary sciences to a multitude of diverse and increasingly distant research fields. It has become an important element in research on computer games or in history to name a few of these domains. At long last it has also arrived in the cognitive sciences where narratives are regarded to be a central means of sense making. From there it was merely a short jump over to the field of cognitive robotics where the concept is employed to describe semantically annotated episodes of recorded activities.</p>
<p>When we observe people or artificial agents performing everyday activities, we collect episodic data that represent trajectories of body parts and activity-specific force events. While these data contain large quantities of information they are, by themselves, meaningless. Only when we put them into a pragmatic context we assign specific meanings to these data. For example, we can interpret the same observed episode as either throwing something or dropping something. This difference in the narrativization, consequently, yields two distinct narratives:</p>
<p>(1) He dropped the glass onto the floor<br />(2) He threw the glass onto the floor</p>
<p>Please note that the situation spawning these two minimal narratives can be identical. We will, therefore need to differentiate between a situation, which has not been narrativized and a description of a situation, which pairs a situation with a selected conceptualization, i.e. interpretation, thereof. In addition to becoming meaningful, this pairing will, in turn, evoke a pragmatic stance that ascribes, for example, a specific perspective and intention to the agent(s) acting in the narrative.</p>
<p>Along with assuming a specific perspective on an episode, narratives also feature a teleological stance and in many cases also a normative valence. This needs to be included to arrive at a comprehensive model that allows for reasoning about narratives, e.g. what the specific differences between two distinct narrativizations of an identical episode are and even what they mean. This type of reasoning would extend the semantics employed, for example, in opinion mining and sentiment analysis, as narratives could then be grouped and compared in terms of perspective, stance or valence.</p>
<p>The contribution of the MUHAI project is to provide a representational framework that can readily be employed in cognitive robotics to counterpart the notion of a task that is given to a robotic agent and can be executed by finding an appropriate action with the notion of a narrative that looks at an action and seeks to makes sense of it. Ideally, one could match the task leading to an action and the narrative describing the action to express if that task has been successfully executed by the agent from the point of view of the narrativizer. As most readers will know there can be vast differences in these judgments, for example, between parents and children concerning the question if a room has been properly cleaned.</p>
<p>Again, it is important to note that the MUAHI approach explicitly rejects an objective notion of a narrative, i.e. to equate a narrative with what has objectively happened and can be recorded a stored as data. As a descriptive notion a narrative assumes a specific point of view on the episodic event. It, therefore, provides a spin on the event and is subjective. Nevertheless, these views can be shared by collectives and become well established frames in which larger historical or everyday episodic events can be seen by societies or groups.</p>
<p>A formal theory of narratives using the logical calculus of OWL-DL, its ontological commitments and underlying foundational framework together with the ontological design patterns, relevant to this work can be found here: <span style="text-decoration: underline;"><a href="http://ceur-ws.org/Vol-2969/paper31-CAOS.pdf" target="blank">http://ceur-ws.org/Vol-2969/paper31-CAOS.pdf.</a></span> The research reported in this paper has been (partially) supported by the FET-Open Project 951846 <code></code>MUHAI -- Meaning and Understanding for Human-centric AI'' funded by the EU Program Horizon.</p>
<p><strong>Credits</strong><br />Photo by <span style="text-decoration: underline;"><a href="https://unsplash.com/@tjump?utm_source=unsplash&amp;utm_medium=referral&amp;utm_content=creditCopyText">Nik Shuliahin</a></span> on <span style="text-decoration: underline;"><a href="https://unsplash.com/s/photos/glass?utm_source=unsplash&amp;utm_medium=referral&amp;utm_content=creditCopyText">Unsplash</a></span></p>]]></description>
			<category>Understanding Everyday Activities</category>
			<pubDate>Fri, 26 Nov 2021 15:20:02 +0000</pubDate>
		</item>
		<item>
			<title>Understanding Everyday Activities</title>
			<link>https://muhai.org/blog/11-understanding-everyday-activities/169-understanding-everyday-activities</link>
			<guid isPermaLink="true">https://muhai.org/blog/11-understanding-everyday-activities/169-understanding-everyday-activities</guid>
			<description><![CDATA[<p><img src="https://muhai.org/images/article/projects_3.png" /></p><h3 style="text-align: left;"><span style="font-size: 14pt;"><a href="https://muhai.org/contact-muhai">Robert Porzel, UHB.</a></span></h3>
<p><img src="https://muhai.org/images/articles/projects_3.png" alt="" /><img src="https://muhai.org/images/article/projects_3.png" alt="" /></p>
<p>If the proof of the pudding is in the eating then the ultimate test for understanding an instruction is its proper execution. This view greatly expands the scope of natural language understanding beyond the usual syntactic and semantic analysis. In this part of the MUHAI project we seek to operationalize the <strong>basic principles of human-centric AI</strong> so that machines will be able to understand how to perform everyday actions in the cooking domain. This involves moving away from executing fully explicit standardised instructions towards understanding instructions conveyed through natural language dialogues. The key challenge here is the integration of world knowledge and pragmatic inferencing into the understanding process, both on the level of language processing and on the level of task execution. For example, the knowledge that chopping a cucumber involves the use of a cutting board and a knife, and presupposes a specific orientation of the cucumber, as well as a conventional slice thickness, is not explicitly mentioned in a recipe, but is essential to carrying out the task and must therefore be inferred from common sense knowledge. Also the build-up of knowledge that generalises across recipes and ingredients is of importance, as it is a precondition for adapting existing recipes to given constraints, and ultimately for the creative design of novel recipes.</p>
<p>In order to achieve these goals we will define two kinds of benchmarks:</p>
<ol>
<li>one that consists in mapping between existing recipes formulated in natural language and actions executed in the VR world</li>
<li>one that allows us to evaluate a new recipe design or variant proposal</li>
</ol>
<p>As in all parts of the MUHAI project the notion of meaning-based and human-centric narratives also applied in the cooking domain. These narratives give meaning to collections of experiences of a virtual agent, i.e. object perceptions, body postures, force dynamics, visual processing and structured data collection, i.e. recipes, images and procedures. Building narratives requires the integration of multimodal sources of input (text, image, sound) and pattern detection in a model of constructional language processing. Constructions will be used as the basic representational unit in which all of these sources are combined. The outcome of constructional language processing is a semantic analysis, including identification of goals, plans, actions, objects, time and causation. The set of analyses make up the starting point for narratives in the domain that can be integrated with the personal dynamic memory in order to truly understand them, in the sense that they can be mapped to a series of low-level actions that can then be executed by a simulated agent in the VR kitchen environment.</p>
<p>To demonstrate the potential of this approach MUHAI will  develop two applications for recipe execution and design:</p>
<ol>
<li><strong>Recipe execution</strong> - This application consists in executing recipes expressed in natural language in a VR kitchen environment. This requires mapping between a recipe (i.e. a sequence of instructions) and a sequence of low-level actions to be executed. The application will involve constructional language processing, consultation with the personal dynamic memory for pragmatic inference, and planning the execution of the concrete cooking actions. The application will be evaluated on the benchmarks described above</li>
<li><strong>Recipe design</strong> - This application is situated in the domain of professional recipe design. In the first part of the project MUHAI will focus on the challenge of building a virtual agent that can act as an assistant chef. This digital assistant needs to integrate technical cooking knowledge with a considerable prior memory of recipes, previously successful and unsuccessful variants, cooking procedures, and cultural context. Most importantly, it needs to do so in an explicable, transparent manner. In a second phase, the focus will shift to a more challenging task that embraces even more aspects of human-centric AI, namely that of recipe design. This is a capacity that goes beyond skill and knowledge and introduces creativity.</li>
</ol>
<p> </p>]]></description>
			<category>Understanding Everyday Activities</category>
			<pubDate>Mon, 08 Mar 2021 09:30:14 +0000</pubDate>
		</item>
	</channel>
</rss>
