WEBVTT

00:00.870 --> 00:03.830
So, welcome to another lecture about organic computing.

00:04.790 --> 00:08.210
Like always, we will have a review of what we had last time.

00:08.910 --> 00:13.750
We talked about CELF-X properties, and we talked about different

00:13.750 --> 00:16.270
aspects of CELF-X properties.

00:16.930 --> 00:21.290
First, we had a very short view what a CELF-X property could be, like

00:21.290 --> 00:24.010
self -healing, optimizing, protecting, configuring.

00:24.370 --> 00:26.190
CELFs are very adaptive and autonomous.

00:26.890 --> 00:33.470
And then, very quickly after that, we talked about multi-agent systems

00:33.470 --> 00:39.250
and this kind of information flow and topology of the inter-agent

00:39.250 --> 00:40.810
relationships.

00:42.230 --> 00:48.550
And what is very important for you to know is coordination mechanisms.

00:48.810 --> 00:53.630
And by coordination mechanisms, I mean that if you have a system, so

00:53.630 --> 00:58.350
like your exercises, you have this car, traffic car, that you are

00:58.350 --> 01:02.570
trying to kind of design to be self-organized.

01:03.150 --> 01:07.410
And this, one thing that's very important is the communication

01:07.410 --> 01:11.050
coordination, how you coordinate these things and what kind of methods

01:11.050 --> 01:12.030
you have for coordination.

01:13.070 --> 01:15.630
So, there are different coordination mechanisms.

01:16.290 --> 01:21.870
We talked about them, and based on a video I showed you about BigDog,

01:21.950 --> 01:26.190
we kind of related that to that video and so on.

01:30.630 --> 01:34.110
And then, we talked about characteristics of CELF-X properties.

01:34.390 --> 01:35.950
There are many characteristics.

01:36.510 --> 01:40.670
Here, we talk only about four characteristics, and we had an

01:40.670 --> 01:41.290
application.

01:41.610 --> 01:44.590
We had the first one was macroscopic or microscopic.

01:44.790 --> 01:49.230
In organic computing, we are interested to have a macroscopic rather

01:49.230 --> 01:50.270
than microscopic.

01:50.650 --> 01:55.710
Macroscopic is that a system behaves as a whole, and then we showed

01:56.290 --> 02:02.210
you an example from immune systems, like in our body, that if a virus

02:02.210 --> 02:06.890
enters the body, all the body is against the virus and protecting

02:06.890 --> 02:07.290
itself.

02:07.790 --> 02:10.990
And the application I showed you, you don't need to go into details of

02:10.990 --> 02:17.270
that application, but it was like a sensor network on a bridge, and

02:17.270 --> 02:20.930
how the sensors, they don't have so much hardware in them, so they

02:20.930 --> 02:24.490
have to kind of, one of them is having lots of memory, the other one

02:24.490 --> 02:28.750
has communication possibilities to others and so on and so on.

02:28.970 --> 02:33.590
And if there is a damage in the bridge, how they can together protect

02:33.590 --> 02:34.890
themselves against that.

02:34.970 --> 02:41.530
That's just an example about macroscopic property.

02:42.010 --> 02:52.290
Now I continue with other property, which is ongoing versus one-shot

02:52.290 --> 02:52.830
property.

02:53.370 --> 02:56.510
What we want is, of course, ongoing property.

02:56.630 --> 03:00.510
When we talk about self-optimization, remember the example about

03:00.510 --> 03:07.350
BigDoc, he must permanently check the self-x characteristic, not just

03:07.350 --> 03:07.930
one -shot.

03:08.050 --> 03:11.530
One-shot means that you have it only once, like you install a

03:11.530 --> 03:13.150
software, it's only once there.

03:13.730 --> 03:19.490
If you always check it, so permanently during the working period you

03:19.490 --> 03:21.610
are checking that, then it will be ongoing.

03:21.870 --> 03:26.690
And in organic computing, we are interested in ongoing properties,

03:26.930 --> 03:32.250
like self-optimizing, that is permanently optimizing, it's working.

03:35.790 --> 03:41.830
Yeah, so here we have ongoing versus one-shot.

03:42.270 --> 03:45.990
And as I told you, one-shot is not that what we are looking for.

03:46.190 --> 03:51.750
We only look for ongoing in organic computing.

03:51.870 --> 03:54.530
An example for that could be self-optimizing.

03:55.930 --> 03:59.270
And the other property is spatial versus non-spatial.

03:59.570 --> 04:01.970
This is also a characteristic of self-x property.

04:02.530 --> 04:06.090
In your exercise, if you have like this traffic jam, it is kind of a

04:06.090 --> 04:06.490
spatial.

04:06.750 --> 04:12.770
Spatial means it depends on the position and the x-y coordinates.

04:12.930 --> 04:16.070
So it's something on a two-dimensional or three-dimensional

04:16.070 --> 04:17.910
coordinate, so it is spatial.

04:18.170 --> 04:22.030
Many applications that we talk about are spatial, but it could be also

04:22.030 --> 04:29.770
non -spatial, like if you talk about colors or properties like sound

04:29.770 --> 04:31.030
properties or whatever.

04:31.410 --> 04:34.210
So that is that it's something that it is not placed in the

04:34.210 --> 04:35.750
environment, it's non-spatial.

04:36.170 --> 04:41.310
In many applications, particularly where we design and measure things

04:41.310 --> 04:46.270
like using entropy and the positions of the cars in the traffic

04:46.270 --> 04:53.750
scenario or learning robots, many applications usually want to have a

04:53.750 --> 04:57.810
kind of a distribution-like property, and this is the spatial

04:57.810 --> 04:58.430
property.

04:59.070 --> 05:04.310
One example for that is if you want to have a specific shape, and this

05:04.310 --> 05:08.170
is like in... the example is in ad hoc network.

05:08.290 --> 05:12.050
Ad hoc network is like a mobile network that we have, the mobile

05:12.050 --> 05:12.450
cells.

05:12.790 --> 05:16.990
So whenever you enter a country, you add your cells to this network.

05:17.370 --> 05:18.650
This is an ad hoc network.

05:19.130 --> 05:21.570
And that should have a certain topology.

05:21.750 --> 05:24.950
So if you have to be in some place, so you have to be in certain

05:24.950 --> 05:33.210
country or a certain area that you enter that cell, that network.

05:33.850 --> 05:38.370
So that should have a certain topology, and that is like a

05:38.370 --> 05:41.230
distribution -like property, a specific shape in there.

05:43.330 --> 05:47.450
The other one was a spatial property is distribution-like.

05:47.530 --> 05:52.570
I already told it, like distribution of a collection of vehicles that

05:52.570 --> 05:58.470
you study in your exercises, or learning robots that are working

05:58.470 --> 06:05.070
together, agents that are producing some specific shape.

06:06.250 --> 06:09.870
So the other characteristic is about resource allocation.

06:10.350 --> 06:15.310
This is also being studied as a characteristic of cell-six property.

06:15.450 --> 06:21.550
We had that actually in a coordination mechanism, so it's like token

06:21.550 --> 06:26.990
-based, or the other one, what was it, tag-based.

06:27.310 --> 06:31.450
You could kind of have a resource allocation mechanism, but this is

06:31.450 --> 06:32.450
also a characteristic.

06:32.830 --> 06:38.650
If a cell-six property is used to have kind of resource allocation,

06:40.270 --> 06:46.570
and resource allocation means that you have some limited amount of

06:46.570 --> 06:51.290
available resources, and you have several entities in your system, and

06:51.290 --> 06:54.450
you want to allocate these resources to your entities.

06:56.370 --> 07:02.850
And resources here, we mean not only something material, but you could

07:02.850 --> 07:06.070
think of a resource in a broader sense.

07:08.050 --> 07:12.750
An example for that, like if there's a limited bandwidth in ad hoc

07:12.750 --> 07:20.590
network, or if you have an area for cleaning robots to clean, everyone

07:20.590 --> 07:23.910
has to have its own partition of this area to clean.

07:24.150 --> 07:27.870
So you kind of have a kind of resource allocation property there.

07:28.110 --> 07:30.290
And that could be in the cell-six property.

07:30.450 --> 07:34.410
So if you have a cleaning robot, so a set of cleaning robots, then

07:34.410 --> 07:41.730
this cell-six property of self-cleaning, for example, would be to have

07:41.730 --> 07:44.030
this kind of resource allocation property in there.

07:45.250 --> 07:49.010
So I was quite quick in this part of the lecture, because it's like

07:49.010 --> 07:49.330
theory.

07:49.450 --> 07:51.230
You can read everything yourself.

07:51.950 --> 07:57.130
But we had these coordination mechanisms and characteristics.

07:57.710 --> 08:02.990
And now how we very shortly want to talk about how you design.

08:03.290 --> 08:08.470
So if you have a set of vehicles, and you want to kind of achieve self

08:08.470 --> 08:11.850
-organization, what are the different steps that you have to go

08:11.850 --> 08:14.830
through to design a self-organizing system?

08:15.510 --> 08:24.170
And here is a picture where we show how you can do that.

08:26.570 --> 08:30.870
First, what you do is that you identify macroscopic properties.

08:31.290 --> 08:38.050
So in organic computing, when we are talking about self-organization,

08:39.270 --> 08:43.070
and we are talking about macroscopic properties.

08:43.470 --> 08:45.330
So you first, you have a system.

08:45.670 --> 08:47.570
Let's say you have a set of cleaning robots.

08:48.190 --> 08:51.650
You have to first, before starting designing that, you have to first

08:51.650 --> 08:56.550
write down what you want to achieve with these robots.

08:56.990 --> 09:00.450
So you write your requirements, a list of whatever you want to have.

09:01.550 --> 09:03.390
They should be macroscopic properties.

09:03.530 --> 09:05.330
What should the system as a whole achieve?

09:05.470 --> 09:07.710
Like they all have to clean a certain room.

09:08.250 --> 09:09.690
This is a macroscopic property.

09:10.670 --> 09:12.650
And then you start the designing.

09:13.930 --> 09:17.310
And for the design, you first need architectural design,

09:17.690 --> 09:18.670
architectures.

09:19.530 --> 09:22.090
We had architectures already in this lecture.

09:22.230 --> 09:23.550
It's a chapter two, I think.

09:23.950 --> 09:27.270
And that's why we have architectures in the beginning of this lecture,

09:27.670 --> 09:30.750
of this course, was that because it's a major step.

09:31.030 --> 09:33.670
If you want to design a system, you need to have an architecture.

09:33.850 --> 09:34.730
What is an architecture?

09:35.030 --> 09:40.950
In architecture, you write about building blocks and the input output

09:40.950 --> 09:44.250
to each block that must exactly be specified.

09:45.190 --> 09:50.030
And there you can write about the coordination mechanisms that we had

09:50.030 --> 09:50.250
here.

09:50.370 --> 09:54.090
So you have your cleaning robots, and then you say, well, coordination

09:54.090 --> 09:59.490
mechanism must be like, we want to have pheromone coordination or tag

09:59.490 --> 10:01.930
-based or whatever, market-based and so on.

10:01.950 --> 10:03.050
So you define it there.

10:04.110 --> 10:06.210
And after that, you have a detailed design.

10:06.330 --> 10:11.150
So you have to specify exactly what these systems, what are the input

10:11.150 --> 10:14.970
outputs to each entity and how they have to behave.

10:15.290 --> 10:16.770
So it's kind of a detailed design.

10:17.510 --> 10:20.970
Then what you do is that you implement that on your system.

10:21.990 --> 10:23.630
So you have tests.

10:23.830 --> 10:25.190
Maybe you do it on a simulation.

10:25.390 --> 10:26.550
You simulate everything.

10:27.510 --> 10:30.930
And then what you have to do is you have to verify that.

10:31.810 --> 10:37.750
So verification is a very important thing, is a very important issue.

10:38.410 --> 10:42.470
If you are working, like imagine those people who work on BigDoc

10:42.470 --> 10:43.210
examples.

10:44.010 --> 10:46.810
So they make BigDoc and they want to sell it to you.

10:47.290 --> 10:48.750
And you say, well, how good is that?

10:48.910 --> 10:49.710
They have to verify.

10:49.850 --> 10:53.770
They have kind of like to give you a sump that it is like in Germany,

10:53.930 --> 10:54.730
you say tooth.

10:55.130 --> 10:58.910
So it must have some certain guarantee that it's working.

10:59.290 --> 11:00.990
It's called quality of service.

11:01.510 --> 11:05.730
So what is, if I have a set of robots cleaning in this room, how is

11:05.730 --> 11:10.130
the quality of service if they are able to clean any room, what kind

11:10.130 --> 11:14.150
of room, what kind of environment, and how do I prove that they are

11:14.150 --> 11:14.810
working well?

11:15.010 --> 11:20.350
So you cannot sell a system without verification mechanism, without

11:20.350 --> 11:23.110
verifying that and showing, okay, this is working.

11:23.830 --> 11:27.590
The problem in organic computing is that it's very complex.

11:28.830 --> 11:34.550
And in the next slide, we show you how we can verify and what are the

11:34.550 --> 11:39.830
tools that you can use to verify certain macroscopic property.

11:40.770 --> 11:43.990
So if you have a microscopic property, it's easy.

11:44.110 --> 11:49.610
So you say, I have one cleaning robot and I can specify, verify the

11:49.610 --> 11:52.010
working of that single robot.

11:52.730 --> 11:56.710
But if you have a macroscopic property, then it is fun.

11:56.930 --> 12:01.510
So you have to really think how the connections are, how the

12:01.510 --> 12:05.310
interactions are and how everything is emerging.

12:05.490 --> 12:09.090
So the emergent effect and so on, you have to all measure these

12:09.090 --> 12:09.370
things.

12:09.630 --> 12:12.390
In the next slide, we talk about verification methods.

12:16.050 --> 12:18.030
So what is verification?

12:19.810 --> 12:25.130
Well, first of all, it is very hard to guarantee about, to give a

12:25.130 --> 12:30.530
guarantee about required self X property, because we don't have a

12:30.530 --> 12:32.270
central or a global controller.

12:32.670 --> 12:37.350
We could have this observer controller as a control entity, but we are

12:37.350 --> 12:38.390
not talking about that.

12:38.630 --> 12:42.150
We are talking about, so when we have a macroscopic, so everything is

12:42.150 --> 12:45.990
connected to each other, it's very difficult that we say, well, we can

12:45.990 --> 12:48.670
kind of control everything from the top.

12:48.770 --> 12:52.790
No, it's sometimes things are happening that we cannot really verify

12:52.790 --> 12:55.210
that from one central point of view.

12:58.250 --> 13:01.530
Well, one thing I showed you before, systems are acceptable in

13:01.530 --> 13:09.670
industrial context, if required behaviors have been verified.

13:10.150 --> 13:15.770
So your future will be in industry or maybe university, but wherever

13:15.770 --> 13:20.610
you work and you build a system, anything that you build, not organic

13:20.610 --> 13:24.790
computing, but anything that you are doing, you have to verify that.

13:25.390 --> 13:29.610
So if you are working in a bank, as the head of the company and you

13:29.610 --> 13:35.570
have a new credit system, you have to verify that or anything else.

13:35.810 --> 13:43.190
So in industrial context, verification is a major step in the system.

13:45.450 --> 13:49.450
Verification methods allow to verify if system has achieved a

13:49.450 --> 13:49.910
requirement.

13:50.110 --> 13:57.230
So you define the requirements in the last slide here and say, we had

13:57.230 --> 14:01.910
this requirement and we see if the system has achieved this

14:01.910 --> 14:02.410
requirement.

14:02.530 --> 14:05.650
And if you forget one part of the requirement, if you forget

14:05.650 --> 14:11.010
something, then you maybe forget to verify that and then you have some

14:11.010 --> 14:12.310
problem in the system.

14:15.540 --> 14:20.800
And the other difficulty is that we have different types of self-fixed

14:20.800 --> 14:21.320
properties.

14:22.040 --> 14:28.020
And for each self-fixed property, you have maybe another verification

14:28.020 --> 14:28.380
method.

14:28.580 --> 14:32.360
So if you are talking about self-configuring, it's different from self

14:32.360 --> 14:33.000
-optimizing.

14:33.180 --> 14:35.640
So if you have self-optimizing, you have to see what kind of

14:35.640 --> 14:39.940
verification methods you use there in self-configuring, self-adaptive

14:39.940 --> 14:40.640
and so on.

14:42.360 --> 14:48.880
And maybe we have to apply multiple verification methods to achieve

14:48.880 --> 14:51.900
all what we have for several self-fixed properties.

14:52.240 --> 14:57.460
So the first verification method is a unit-based verification.

14:58.340 --> 15:05.600
And it's like if you have a set of robots, you test each robot alone,

15:05.940 --> 15:07.060
say, well, this is working.

15:07.620 --> 15:10.700
And then you do another unit and you test that there.

15:11.060 --> 15:14.580
And you test each unit in your system separately.

15:14.760 --> 15:18.520
This is a unit-based verification method.

15:20.420 --> 15:21.360
Isolate individual robots.

15:21.360 --> 15:23.080
So you isolate individual parts and see if they are correct.

15:25.200 --> 15:28.740
And the other measure is integration testing.

15:29.800 --> 15:34.780
And it's like you have those systems that you verified individually.

15:35.200 --> 15:38.760
And then you combine them as a group and see if they are working.

15:40.580 --> 15:45.060
If you have microscopic self-fixed property, forget the unit-based.

15:45.740 --> 15:51.900
Maybe you can test the system single entities if they are working or

15:51.900 --> 15:52.040
not.

15:52.100 --> 15:53.080
And you must do that.

15:53.440 --> 15:58.940
But you cannot verify a macroscopic self-fixed property using the unit

15:58.940 --> 15:59.280
-based.

15:59.920 --> 16:02.220
So you can do it microscopic level.

16:02.540 --> 16:07.620
But for macroscopic system as a whole, you cannot use unit-based

16:07.620 --> 16:08.260
methods.

16:09.580 --> 16:12.600
The other one is the integration testing.

16:13.220 --> 16:20.360
And also these unit tests, they could be useful for one-shot property.

16:20.560 --> 16:24.280
So one-shot property, we said in the beginning of the lecture today.

16:24.580 --> 16:27.780
One-shot property is that a property that you have it only once in the

16:27.780 --> 16:28.060
system.

16:28.160 --> 16:32.960
Like you install something or it's only for one time.

16:33.440 --> 16:35.780
Then you could use this.

16:35.900 --> 16:38.180
You say, well, I verify everything and so on.

16:38.200 --> 16:38.840
And it works.

16:38.980 --> 16:43.020
But for ongoing property, that is our goal, they cannot be used.

16:44.440 --> 16:49.060
So, conclusion here, this unit-based and integration, they are good

16:49.060 --> 16:50.380
for...

16:50.380 --> 16:55.620
might be good for microscopic properties, but not definitely

16:55.620 --> 16:56.980
macroscopic.

16:59.040 --> 17:00.960
Then we have formal proof.

17:01.180 --> 17:04.660
So if you are a good mathematician, you can sit down and write a

17:04.660 --> 17:06.600
mathematical proof of your system.

17:07.280 --> 17:11.760
You can write it down how you prove the functionality of your system.

17:11.860 --> 17:14.580
You have to model it very well and then prove it.

17:17.270 --> 17:21.460
The point is, a formal modeling of such system is very difficult.

17:21.720 --> 17:25.840
How you want to mathematically model this complex system, it's very

17:25.840 --> 17:30.140
difficult and it's almost impossible to do that.

17:32.250 --> 17:36.600
It could be a formal proof, maybe feasible for simple properties,

17:37.160 --> 17:39.260
like, again, microscopic properties.

17:40.920 --> 17:45.740
So for a microscopic property, like for a single robot or a single

17:45.740 --> 17:49.800
car, you can write it, okay, what kind of functionality you want.

17:50.880 --> 17:54.560
For example, self-optimizing, you can write some formula and prove it

17:54.560 --> 17:57.280
and say, well, this must work according to this formula.

17:58.320 --> 18:02.520
But for a macroscopic where you see it from the top, it's very

18:02.520 --> 18:05.040
difficult and cannot actually be used.

18:06.040 --> 18:11.280
The method that we actually use, it's statistical experimental

18:11.280 --> 18:11.580
verification.

18:12.520 --> 18:13.460
What does that mean?

18:13.760 --> 18:18.840
It means that you make, you have a system of these different entities

18:18.840 --> 18:26.600
and you run it 100 times, 300 times, million times, as much as you

18:26.600 --> 18:31.300
can, the different parameter settings and see if the system is always

18:31.300 --> 18:31.740
working.

18:32.060 --> 18:36.060
So the more you experiment on that system, the more you are sure that

18:36.060 --> 18:37.280
it's working well.

18:37.720 --> 18:43.340
And then in the quality of service at the end of the delivering to the

18:43.340 --> 18:48.900
industrial partners, you can say, well, you have run it for a million

18:48.900 --> 18:52.120
times with different parameter settings and it worked.

18:52.620 --> 18:57.140
So this is the proof, like, and then they have to believe you or not.

18:57.560 --> 19:02.800
So for giving a good verification, experimental verification, you have

19:02.800 --> 19:05.900
to have a large sample sufficient.

19:06.160 --> 19:09.840
So the sample must be sufficiently large.

19:10.800 --> 19:15.760
And also if you have several parameters, the parameter space must be

19:15.760 --> 19:16.600
covered well.

19:17.240 --> 19:24.860
So it means if I take the example about the big dog example, so that

19:24.860 --> 19:30.200
was the system which is walking, like it must be able to work in any

19:30.200 --> 19:30.720
situation.

19:31.380 --> 19:33.180
So now you change the parameters.

19:33.500 --> 19:38.600
First, you go to the snow area, then you go to the stony area, then

19:38.600 --> 19:42.000
you go to the normal area, then you change it for all the parameter

19:42.000 --> 19:45.660
settings that you could give into the system and see if that system is

19:45.660 --> 19:46.000
working.

19:47.540 --> 19:54.220
And then you have to say that the system, so then you give a kind of

19:54.220 --> 20:00.960
verification stamp on the system and say, well, that is, that's

20:00.960 --> 20:02.760
working for that situation.

20:03.780 --> 20:08.580
It's very expensive if you have a system like the big dog or even a

20:08.580 --> 20:09.080
simulation.

20:09.400 --> 20:15.560
If you run it, let's say for 20 parameters, 20 combinations, that is

20:15.560 --> 20:24.200
really low, and a million times experiments, 20 parameters, you can

20:24.200 --> 20:26.900
see how much, how many experiments you have to run.

20:27.340 --> 20:28.540
Imagine your exercises.

20:28.760 --> 20:32.840
If you want to really verify the functionality of your system with all

20:32.840 --> 20:36.780
the input variables, then you have to run every input variable, which

20:36.780 --> 20:42.420
is a random variable for at least 100 times, and then make average

20:42.420 --> 20:45.020
over that, and then change parameters and so on.

20:45.080 --> 20:50.260
This is really, really time-consuming, and it is very expensive if the

20:50.260 --> 20:51.840
simulations are taking time.

20:52.140 --> 20:55.940
Sometimes the simulations take like three hours just running one

20:55.940 --> 21:01.920
thing, or even for chemical engineering, you run like five days just

21:01.920 --> 21:02.660
one experiment.

21:03.200 --> 21:06.380
Now you have to run it for five million times, then you can imagine

21:06.380 --> 21:08.880
how many computation power you need.

21:09.080 --> 21:12.580
Even if you paralyze that, then still it's very expensive.

21:15.560 --> 21:23.060
So you can also use some metrics for verification methods.

21:23.240 --> 21:28.360
So many of the verification methods require a metric that reflects

21:28.360 --> 21:30.740
this self-x property.

21:31.380 --> 21:39.480
You learn entropy, but this is not only the only way of doing... it's

21:39.480 --> 21:42.880
not the only way of measuring self-x property.

21:43.600 --> 21:47.580
But you can use that, for example, for some spatial distributions.

21:48.060 --> 21:52.860
You can use this entropy value, or even you remember the example with

21:52.860 --> 21:56.360
the colors, different green and red colors.

21:56.400 --> 21:59.380
You could measure the order in the system.

21:59.540 --> 22:01.020
You could use this entropy.

22:01.560 --> 22:05.620
And entropy is a measurement that you can have to measure macroscopic

22:05.620 --> 22:06.140
property.

22:08.640 --> 22:11.300
The other metric is the distance.

22:11.580 --> 22:15.640
So what you do is that you do this experiment, and every time that you

22:15.640 --> 22:21.500
do experiments, you say you see how far you are from the optimal point

22:21.500 --> 22:24.140
that you already require in your system.

22:24.200 --> 22:28.420
So every time that you run an experiment, you either measure entropy,

22:28.660 --> 22:31.420
or you measure the distance to the ideal point.

22:32.240 --> 22:34.880
The system must have an ideal point, and you measure it.

22:34.980 --> 22:39.580
Like the big dog, if you kick it, you measure how far it goes far from

22:39.580 --> 22:43.180
the ideal, what you measure you must have in your system.

22:44.260 --> 22:48.160
This is actually a measure that always has been used, also in

22:48.160 --> 22:49.320
optimization area.

22:49.680 --> 22:55.560
If you look at optimization problems, those who had the lecture about

22:55.560 --> 22:59.380
optimization here, you know every time you want to measure how good an

22:59.380 --> 23:02.420
algorithm is, you measure it, the distance to the global optimal

23:02.420 --> 23:04.620
solution, if it is known.

23:07.400 --> 23:12.580
And this metric, the distance is that you have to be able to measure

23:12.580 --> 23:16.760
it at each moment, so that you have to see how you measure it at each

23:16.760 --> 23:20.460
time step, because we are talking about ongoing property in a self

23:20.460 --> 23:21.060
-fixed property.

23:23.940 --> 23:28.560
And, well, other metrics are statistical values like variance and

23:28.560 --> 23:29.060
average.

23:29.440 --> 23:34.380
So, you either measure entropy, or distance, or both, and then you

23:34.380 --> 23:38.600
can, you run the experiments 10,000 times, a million times, and every

23:38.600 --> 23:41.880
time you have to make an average, and you say, you give a standard

23:41.880 --> 23:46.720
deviation, or a standard error, what is the error actually in your

23:46.720 --> 23:47.180
system?

23:47.960 --> 23:50.780
How worse can your system be?

23:52.440 --> 23:52.780
Good.

23:53.140 --> 23:55.200
So, that was all about verification.

23:56.240 --> 24:01.440
There are courses about verification methods, and model checking, or

24:01.440 --> 24:05.360
formal verifications, you can go, I think, to the computer science

24:05.360 --> 24:07.540
department here and visit courses about that.

24:08.080 --> 24:11.780
It's very mathematics, so you have lots of mathematics in there, and

24:11.780 --> 24:15.700
if you are interested to see what kind of tools there are, you could

24:15.700 --> 24:16.820
go and visit those courses.

24:17.140 --> 24:20.300
We just stop here and we say, well, these are the ways that you can

24:20.300 --> 24:26.580
use to verify, and we will, as in your exercises, we will stick to the

24:26.580 --> 24:29.660
experiments and these metrics that I showed you.

24:30.520 --> 24:37.000
Now, I want to give some examples, not to get so dry in the theory.

24:38.580 --> 24:41.380
The first example is self-healing.

24:43.820 --> 24:48.060
So, what is self-healing, and what kind of properties we need for self

24:48.060 --> 24:48.300
-healing?

24:48.700 --> 24:52.800
So, if a system is self-healing, first of all, it has to know that

24:52.800 --> 24:53.680
there was a problem.

24:54.560 --> 25:00.700
So it has to have the property to detect the problem, to detect that

25:00.700 --> 25:06.640
something is wrong in the system, and then he has to kind of know how

25:06.640 --> 25:08.600
to recover to the normal situation.

25:09.120 --> 25:12.420
So, the properties are listed here.

25:12.880 --> 25:18.300
Detect improper operations, initiate corrective actions before even

25:18.300 --> 25:18.940
they occur.

25:19.660 --> 25:24.980
This is critical, so they have to have some kind of prediction

25:24.980 --> 25:27.260
mechanisms or something like that.

25:31.520 --> 25:38.280
And even, so as we also wrote here, before they occur or even prevent

25:38.280 --> 25:43.800
failures by potential attacks or damages, that goes back to the other

25:43.800 --> 25:44.160
point.

25:44.540 --> 25:45.680
And then also learning.

25:47.160 --> 25:52.000
So, learning is a very important aspect that we also start in the

25:52.000 --> 25:54.000
lecture five about that today, maybe.

25:55.480 --> 26:00.520
But then he has to know, so in the self-healing, it�s like if there is

26:00.520 --> 26:04.640
a damage in the system and he wants to, he has to know what kind of

26:04.640 --> 26:07.540
damage, so he has to learn it in the system, what kind of damage can

26:07.540 --> 26:10.940
happen and how can he predict that that damage could happen.

26:11.180 --> 26:12.200
So he has to learn.

26:12.840 --> 26:20.320
And for learning, he can have experiences or a knowledge pool, and he

26:20.320 --> 26:23.060
uses this knowledge pool for future decisions.

26:23.600 --> 26:31.500
You will see in the chapter about machine learning how we use this

26:31.500 --> 26:33.360
knowledge pool for learning.

26:34.740 --> 26:42.180
So, self-healing systems, they are very useful in situations where we

26:42.180 --> 26:43.680
need continuous services.

26:44.140 --> 26:49.060
So the service must permanently be there, and we cannot wait until a

26:49.060 --> 26:52.120
human being goes there and provides the service.

26:52.900 --> 26:57.400
So we need something, so this should be working ongoing in the system,

26:57.900 --> 27:00.760
and manual repair is not feasible.

27:00.760 --> 27:06.020
You will see in an example, I think, well, one example is in an ad hoc

27:06.020 --> 27:06.520
network.

27:06.860 --> 27:10.360
So these are the mobile cell connected to each other, and now if a

27:10.360 --> 27:17.060
mobile cell is out, what happens is that these two must get connected

27:17.060 --> 27:20.980
to each other, these two nodes, and they must find a way that they

27:20.980 --> 27:23.880
say, okay, they don�t have any connection to each other.

27:24.160 --> 27:29.120
They have to see how they get the connection if their interconnecting

27:29.120 --> 27:34.860
agent is gone, so they kind of find the next neighbor to connect to

27:34.860 --> 27:35.220
each other.

27:35.420 --> 27:40.040
So they have to organize themselves and adapt continuously to the

27:40.040 --> 27:41.080
changes in the network.

27:41.660 --> 27:46.100
This is a simple example and just very abstract, but another

27:46.100 --> 27:51.920
interesting application is if you have autonomous mobile robots in

27:51.920 --> 27:52.900
human environments.

27:53.260 --> 27:58.640
So imagine that we have a robot here, and it has to walk like Big Dog,

27:58.740 --> 28:00.520
but not that intelligent like Big Dog.

28:01.260 --> 28:04.880
So what happens is that we enter the robot, come actually, I�ve seen

28:04.880 --> 28:06.880
that robot that can enter this room.

28:07.740 --> 28:10.060
The point is that environment changes, right?

28:10.160 --> 28:13.860
So one day a table is here, the other day the table is not here, or

28:13.860 --> 28:16.300
chairs or people are sitting, they can move.

28:16.740 --> 28:18.500
The environment can dynamically change.

28:19.500 --> 28:22.960
So we don�t have any model of the environment, so we cannot train the

28:22.960 --> 28:26.920
system to the environment, and we don�t know what kind of mistakes can

28:26.920 --> 28:27.300
happen.

28:28.820 --> 28:34.080
And then the project that will come in the next slide, they had two

28:34.080 --> 28:34.500
goals.

28:34.860 --> 28:40.240
They wanted to be safe, it means that their robot works still, and

28:40.240 --> 28:41.740
also fault tolerant.

28:41.900 --> 28:45.280
So if there is a fault in the system, still it has to walk.

28:48.240 --> 28:51.980
So what we have, uncertainties, unknown environment, and unforeseen

28:51.980 --> 28:56.660
situations, so you cannot predict a failure that is not happening.

28:57.020 --> 29:01.140
So if you know that there is a failure, then you can kind of work on

29:01.140 --> 29:02.640
that, and it won�t never happen.

29:02.920 --> 29:05.860
So a failure is something that you can never know that this is a

29:05.860 --> 29:08.540
failure, otherwise you can�t deal with that before it happens.

29:11.940 --> 29:16.720
So for this robot that is walking, they said, okay, they want to

29:16.720 --> 29:20.340
detect, react, and adapt to malfunction.

29:20.720 --> 29:24.640
So if there is something wrong with the robot, it has to detect it and

29:24.640 --> 29:26.320
do it automatically.

29:28.020 --> 29:31.260
The system must work, no critical system status.

29:31.800 --> 29:35.960
It must always work, and low cost implementation.

29:36.460 --> 29:37.920
That was this critical point.

29:37.980 --> 29:42.980
They didn�t want to have a very difficult program to deal with every

29:42.980 --> 29:48.420
failure, and they came up with a very interesting application.

29:49.880 --> 29:52.840
And this is this robot called Oscar.

29:53.700 --> 30:00.580
It�s like a spider, and people in Zubek, Professor Mele, he worked

30:00.580 --> 30:05.500
with biologists, and the biologists, what they did, I mean, we had

30:05.500 --> 30:10.640
actually the videos, but we don�t show you, but they took a spider and

30:10.640 --> 30:15.540
they cut off the legs, one by one, and they let it run, this spider,

30:15.740 --> 30:18.900
and they analyzed biologically how this spider is working.

30:20.100 --> 30:26.820
So what happens is that if you have a spider or an insect, never do

30:26.820 --> 30:28.880
it, but the biologists, they did that.

30:29.580 --> 30:34.100
So how people are walking, actually, have you ever seen how you

30:34.100 --> 30:34.940
yourself walk?

30:35.800 --> 30:40.240
The way you walk is that you have one leg standing on the floor, and

30:40.240 --> 30:42.260
the other leg then can swing.

30:46.960 --> 30:49.280
And then you have stance, and then swing.

30:50.040 --> 30:51.820
If you are a spider, what happens?

30:51.940 --> 30:53.320
The same thing happens in there.

30:53.680 --> 30:58.920
So you at least need to have one leg standing, the next neighboring

30:58.920 --> 31:02.620
leg can then swing, and so on, right?

31:03.100 --> 31:05.460
So they test that on a spider.

31:05.980 --> 31:08.220
I think spiders have eight legs, I don�t know.

31:08.560 --> 31:13.300
Well, they started cutting the legs, and they realized that with three

31:13.300 --> 31:15.840
legs they can still walk, and very quickly.

31:17.640 --> 31:23.300
So they programmed that on this Oscar, and what happens here, these

31:23.300 --> 31:27.780
legs, here you see there is a joint, and this joint, they can

31:27.780 --> 31:31.060
automatically do this amputation property.

31:31.260 --> 31:35.080
So at some point they have a remote controller, say a leg number, for

31:35.080 --> 31:36.780
example, three, can amputate.

31:36.860 --> 31:42.820
And what happens is that this lower part can fold back to here, so it

31:42.820 --> 31:47.360
means that this leg is amputated, so kind of, so it�s not on the

31:47.360 --> 31:47.600
floor.

31:48.840 --> 31:52.180
So they first wrote a program which is very simple.

31:52.280 --> 31:55.460
So the goal was that we want to have something simple, but if

31:55.460 --> 31:58.580
something happens, like a malfunction happens, we want to have a self

31:58.580 --> 32:01.160
-healing property, so the system still works.

32:02.320 --> 32:09.020
So they said, okay, every leg is an agent, and every agent has a

32:09.020 --> 32:13.700
simple program, so it has a rule for leg I.

32:14.000 --> 32:20.940
You have a rule, and it says, okay, leg I, if the left-hand side leg

32:20.940 --> 32:28.100
of this leg has a ground contact, and the right-hand side has a ground

32:28.100 --> 32:36.000
contact, so if both neighbors of certain leg I have a ground contact,

32:36.400 --> 32:37.200
you can swing.

32:39.060 --> 32:40.780
So this is so, you go to the swing.

32:41.560 --> 32:46.100
If they don�t have, one of them doesn�t have a ground contact, it

32:46.100 --> 32:48.680
means that you cannot swing, so you have to stand.

32:49.540 --> 32:50.960
And this is this stance.

32:52.640 --> 32:55.600
So, and then they also defined, well, it�s not.

32:55.980 --> 32:59.700
It�s the main program, but this is a very robust program.

32:59.840 --> 33:03.220
You will see in the next slide, if you amputate a leg, it�s still

33:03.220 --> 33:05.260
working because neighbors are neighbors.

33:05.380 --> 33:05.940
It doesn�t matter.

33:06.100 --> 33:09.480
So at every time you look at the neighbors, if one neighbor doesn�t

33:09.480 --> 33:13.420
exist, you look at the other one that is still your neighbor, and then

33:13.420 --> 33:15.820
you can check that neighbor.

33:16.800 --> 33:23.240
So swing phase has constant length in this example, and the duration,

33:23.440 --> 33:29.560
how long a leg is standing, it defines the velocity, right?

33:29.600 --> 33:34.100
So if I stand here and I have one leg swinging, still I have a

33:34.100 --> 33:38.300
standing leg, I have a velocity of almost zero, right?

33:38.460 --> 33:43.640
So you can define how long you stand, and then you can find a

33:43.640 --> 33:44.020
velocity.

33:46.060 --> 33:51.580
And this is actually, the legs, they have kind of coordinated, so you

33:51.580 --> 33:58.400
have the two in front, front left, front right, middle right, middle

33:58.400 --> 34:04.240
left, and back, so, well, HR and HL.

34:04.240 --> 34:08.620
So these are the orientations.

34:09.240 --> 34:13.220
So now I show you the video, how it is walking.

34:20.140 --> 34:22.440
So do you see it?

34:24.100 --> 34:28.860
So this is, so if you look at one certain leg, you see that it walks,

34:29.540 --> 34:32.740
it swings when the neighbor is standing.

34:34.260 --> 34:35.520
And so this still works.

34:35.900 --> 34:42.420
So this Oscar is as big as, well, quite big, so you can see in the

34:42.420 --> 34:46.440
conferences they bring that as a demo, and you see that it is walking

34:46.440 --> 34:47.120
in the system.

34:47.400 --> 34:47.640
Yes?

34:49.660 --> 34:53.560
Well, they turn it on, it's like a robot.

34:57.500 --> 34:59.540
Well, the front legs are starting.

34:59.860 --> 35:03.920
Well, details, I have to go through the details, how it is actually

35:03.920 --> 35:06.960
being programmed, but let me run it again.

35:09.380 --> 35:09.860
No.

35:12.160 --> 35:12.920
Oh, yeah.

35:14.540 --> 35:21.560
So one leg will, at least, I think two legs must start, two not

35:21.560 --> 35:26.960
neighboring legs must start, one after the other, maybe, and then that

35:26.960 --> 35:28.100
will start working.

35:31.880 --> 35:34.000
I have to go through the details, that's a good question.

35:34.100 --> 35:36.780
So the question for those who are listening to the recording is that

35:36.780 --> 35:38.340
how the system starts walking.

35:38.760 --> 35:43.580
Well, it starts walking, I think, then you have to define which leg is

35:43.580 --> 35:55.200
the front leg, and which leg is initiating the walk, and definitely

35:55.200 --> 36:02.480
one of them will start, and then the others, at least two must swing,

36:02.580 --> 36:04.160
right, that the system works.

36:04.840 --> 36:06.180
That we have to check.

36:06.520 --> 36:09.940
I think two or three must swing that the system works, right?

36:10.960 --> 36:11.640
How is it?

36:11.920 --> 36:12.600
Two or three?

36:14.300 --> 36:15.100
Two or three?

36:17.260 --> 36:22.320
Yeah, well, let's look at the spider, or I check the document.

36:23.140 --> 36:23.340
Okay.

36:24.280 --> 36:30.540
So now what we do is that we can, so if one leg is amputated, what

36:30.540 --> 36:34.640
happens is that, so it is an ongoing property, imagine that the spider

36:34.640 --> 36:40.900
is walking, and now they kind of, with a remote control, they amputate

36:40.900 --> 36:44.400
one leg, and it means that one leg is clapped back, so it's like

36:44.400 --> 36:45.320
folded back.

36:45.940 --> 36:50.800
And what happens is that because we have a single rule for, and then

36:50.800 --> 36:57.580
we have a change in the neighborhood relation, nothing happens.

36:57.700 --> 36:59.220
The system is very robust.

36:59.580 --> 37:04.180
So every leg looks at, so if the neighbor is gone, it looks at the new

37:04.180 --> 37:08.240
neighbor, and then as long as you have enough new neighbors, you can

37:08.240 --> 37:09.700
kind of walk, right?

37:09.740 --> 37:14.140
So it's a system, it's self-healing in a way that it can detect, okay,

37:14.180 --> 37:17.140
a neighbor is gone, what is the next neighbor, and the algorithm is

37:17.140 --> 37:20.920
the same as before, so you don't need to change it.

37:22.360 --> 37:26.820
So the other property, self-x property, is self-optimizing or self

37:26.820 --> 37:27.560
-optimization.

37:28.020 --> 37:35.120
So a system is self-optimizing if it's permanently or always tries to

37:35.120 --> 37:36.400
optimize it's working.

37:36.880 --> 37:39.540
What properties do we need for self-optimizing?

37:39.880 --> 37:42.800
First of all, we have to analyze the current situation.

37:43.320 --> 37:46.840
So we have to see if it is in optimal state or not.

37:47.140 --> 37:51.620
For that, you need an observer or observation mechanism like some

37:51.620 --> 37:55.580
sensors or monitoring systems, whatever, that you say what is the

37:55.580 --> 37:56.280
current status.

37:57.260 --> 37:59.840
And then you need some kind of optimization.

38:01.260 --> 38:05.260
The point is, if you have a difficult problem, difficult system, and

38:05.260 --> 38:11.380
the optimization takes like two days to optimize, then it's not good,

38:11.400 --> 38:14.460
because it takes time to optimize now for the new situation.

38:15.120 --> 38:20.100
So for that, we have to have a selection, so a list of optimal

38:20.100 --> 38:22.500
solutions somewhere in the memory.

38:22.980 --> 38:27.140
And then for the new change, for the changed situation, we can select

38:27.140 --> 38:28.580
one of these solutions.

38:29.720 --> 38:38.640
So that's why we have here three options for optimization, because it

38:38.640 --> 38:40.280
could be that optimization takes long.

38:40.380 --> 38:44.200
So if optimization, so generating a new optimal solution takes long,

38:44.560 --> 38:49.400
then we can say, well, we can make an offline optimization somewhere

38:49.400 --> 38:51.260
else on a simulation or before.

38:51.900 --> 38:56.680
And then we can select one of the optimal solutions for the current

38:56.680 --> 38:58.360
status from a list.

38:59.180 --> 39:02.580
That's why it is right here, offline optimization.

39:04.000 --> 39:09.300
The other thing is that we have kind of... the parameter is a little

39:09.300 --> 39:11.800
bit different from the optimal.

39:12.860 --> 39:17.140
So you must, from the optimal solution, so you must not run a new

39:17.140 --> 39:21.880
optimization algorithm with a new, everything randomized input

39:21.880 --> 39:26.920
variables, but you can kind of select that, take that solution that is

39:26.920 --> 39:27.280
there.

39:27.820 --> 39:31.780
You say, okay, it's not so far from the optimal solution.

39:31.940 --> 39:35.660
You try to change it to some kind of local optimization to change it

39:35.660 --> 39:37.760
online to adapt to the new situation.

39:38.720 --> 39:40.160
So this is the other way.

39:40.700 --> 39:43.920
And the other option is that you generate new solutions.

39:44.240 --> 39:47.380
So you write an efficient optimization algorithm.

39:47.920 --> 39:52.100
And this optimization runs very quickly because the problem is not so

39:52.100 --> 39:53.420
difficult to solve.

39:53.640 --> 39:59.560
So every time that you need an optimization, you run an optimization

39:59.560 --> 40:00.180
algorithm.

40:01.260 --> 40:02.600
So it depends on the problem.

40:02.680 --> 40:05.060
That's why we have three options here.

40:06.960 --> 40:09.700
After that, you optimize.

40:09.960 --> 40:14.860
You have then to assign these variables to the goals, to adapt the

40:14.860 --> 40:16.300
system behavior to the goals.

40:18.620 --> 40:23.080
The other aspect in optimization is you can also have a learning

40:23.080 --> 40:29.680
there, particularly in offline optimization or the so-called online

40:29.680 --> 40:30.240
optimization.

40:31.020 --> 40:35.880
So you will see today, I think, you will see that how we can have

40:35.880 --> 40:43.560
optimization and how can we learn to have all these optimal parameters

40:43.560 --> 40:47.160
and how can we learn from experiences in knowledge pool.

40:48.580 --> 40:51.540
The other property we have is a self-protection.

40:51.900 --> 40:56.420
So until we had self-healing, self-optimizing, now self-protection.

40:57.240 --> 41:01.620
The system has to protect itself against the damages and all the

41:01.620 --> 41:03.000
malfunctions that can happen.

41:03.120 --> 41:07.260
So if I tell you self-protection, always remember virus scanner.

41:07.500 --> 41:13.400
Well, it's a very abstract way of saying that, but if you want to have

41:13.400 --> 41:18.900
some kind of a keyword to think of, protecting means firewall and so

41:18.900 --> 41:20.960
on, but not in the form that you think.

41:21.080 --> 41:25.240
In self, its property must be more intelligent than a firewall or

41:25.240 --> 41:26.100
something like that.

41:26.800 --> 41:33.560
So it has to detect and identify hostile behaviors and has to take

41:33.560 --> 41:34.700
autonomous actions.

41:38.380 --> 41:45.060
And that's the example I told you about the virus scanner or firewalls

41:45.060 --> 41:45.600
and so on.

41:46.080 --> 41:50.480
These are that we have for many years in the system.

41:50.860 --> 41:54.780
So it's not self-protection, it's only protection mechanism and one

41:54.780 --> 41:56.060
-shot protection mechanism.

41:56.380 --> 41:59.920
But if you want to have self-protection, then you have to have a

41:59.920 --> 42:02.340
system that does it itself.

42:02.680 --> 42:06.520
So a virus scanner or virus protection or firewall is a system that

42:06.520 --> 42:08.180
you install as a human being.

42:08.620 --> 42:12.720
But a self-protecting system, imagine a system like a cleaning robot

42:12.720 --> 42:19.220
or a big dog, they must protect themselves, not by a human user going

42:19.220 --> 42:22.600
there and reinstalling and checking what is to be done.

42:23.060 --> 42:28.120
But you have to... they have to do it themselves.

42:28.600 --> 42:35.300
That's why what we have here for many years, it's not definitely self

42:35.300 --> 42:37.120
-protecting, it's a protection mechanism.

42:37.160 --> 42:42.220
These are protection mechanisms, but are not self-protecting.

42:42.520 --> 42:42.920
And why?

42:43.060 --> 42:47.460
Because the need for that, for these systems and IT personnel to

42:47.460 --> 42:49.520
manually examine and analyze.

42:49.520 --> 42:55.580
And if you have a human being, then it can be that there are some

42:55.580 --> 42:58.200
errors from the human being.

42:58.560 --> 43:04.200
So the human being is not always the best way to... sometimes you see

43:04.200 --> 43:07.040
the supervisors, they install something and there is a mistake and

43:07.040 --> 43:08.820
then you have to call them back and so on.

43:09.080 --> 43:13.800
So if the system is self-protecting, it might be even much better than

43:13.800 --> 43:20.400
having a human being working there because they can protect themselves

43:20.400 --> 43:21.920
against a human.

43:26.180 --> 43:26.420
Yeah.

43:27.980 --> 43:32.040
So for all these systems, we need learning for all these properties.

43:32.180 --> 43:36.820
So it means that here you can also have a learning mechanism.

43:37.740 --> 43:42.500
When you want to detect and identify behaviors, you have to know, you

43:42.500 --> 43:43.380
can learn that.

43:47.840 --> 43:49.260
Yeah, well, that could be a way.

43:51.300 --> 43:57.840
Well, that could be... so here we don't exactly define what is... so

43:57.840 --> 44:03.380
in the specification here, we don't exactly define what we mean with

44:03.380 --> 44:04.760
protection in a way.

44:04.840 --> 44:10.940
But you can say, well, the system has a list of different options or

44:10.940 --> 44:12.000
its requirements.

44:12.200 --> 44:15.340
So it knows that it must work like this and that and this.

44:15.760 --> 44:23.260
And if there is another... some hostile behavior entering, so entity

44:23.260 --> 44:27.240
entering the system, it kind of has to protect itself.

44:28.020 --> 44:32.880
Like a virus, when a virus enters a system, IT infrastructure, it has

44:32.880 --> 44:35.520
to say, well, okay, now I have a virus, now I stopped working.

44:36.160 --> 44:40.520
In a normal life that we have, so it is some... we get a message,

44:40.660 --> 44:44.100
there is a virus, do I do something there or not?

44:44.140 --> 44:45.180
And then we have to click.

44:45.560 --> 44:48.440
But in an autonomous system, they have to do it themselves.

44:48.680 --> 44:53.620
They say, well, it is not in my list, as you say, then I put it.

44:53.660 --> 44:55.280
And then they can also learn that.

44:56.640 --> 45:01.580
But we don't write here about learning because learning is... it is

45:01.580 --> 45:07.660
difficult to learn for the self-protection because learning is not

45:07.660 --> 45:08.540
always...

45:08.540 --> 45:11.300
we cannot achieve a hundred percent of learning.

45:11.860 --> 45:15.160
And I tell you in the next slides how we can achieve learning.

45:15.440 --> 45:22.200
So we are not now in the machine learning area and cognitive systems,

45:22.260 --> 45:23.400
we are not that far.

45:24.500 --> 45:27.880
We can learn very simple things.

45:28.820 --> 45:32.860
We can now... for example, if you look at the games that are being

45:32.860 --> 45:38.240
learned, in the 90s we learned Deep Blue.

45:38.540 --> 45:44.940
Deep Blue is a computer who plays chess and that was not still... that

45:44.940 --> 45:48.480
was not learning, a hundred percent learning.

45:49.210 --> 45:53.840
But we can learn some parameters for the system.

46:01.570 --> 46:05.690
So these are three aspects again.

46:06.590 --> 46:10.210
They can scan for suspicious activities and react.

46:10.550 --> 46:13.730
So that was what we had for several years.

46:14.310 --> 46:17.270
And reacting is something new in a self-protection system.

46:17.910 --> 46:21.890
And they also need to provide information to the person, to the

46:21.890 --> 46:24.210
responsible person in the system.

46:24.830 --> 46:30.350
And what the property is that they can avoid human errors that are

46:30.350 --> 46:32.110
happening by mistakes in the system.

46:32.970 --> 46:36.230
So the next property is self-configuring.

46:36.690 --> 46:39.290
So a system that is configuring its property.

46:40.790 --> 46:45.630
And it's because we wanted... because self-configuring systems

46:45.630 --> 46:46.630
dynamically...

46:46.630 --> 46:51.070
adapt dynamically to changes in the IT environment with little human

46:51.070 --> 46:51.890
intervention.

46:52.570 --> 46:54.550
This is the goal of self-configuration.

46:56.130 --> 47:02.810
And it means that it usually works for a system... for example, if you

47:02.810 --> 47:08.910
have a set of sensors somewhere, you throw them somewhere and then you

47:08.910 --> 47:10.510
want to reconfigure them.

47:11.070 --> 47:12.510
So you can't do it manually.

47:12.650 --> 47:16.230
You have to find the sensors and then run the program again and so on.

47:16.230 --> 47:20.530
But if the system can kind of adapt itself to the new environment and

47:20.530 --> 47:27.010
reconfigure itself, then for some applications it's more practical.

47:27.990 --> 47:32.630
But it must know several... it must have several properties.

47:35.410 --> 47:39.710
First of all, they have to have the ability to know how to run the

47:39.710 --> 47:40.070
program.

47:40.490 --> 47:44.090
So if you reconfigure a system like install a new Windows system, for

47:44.090 --> 47:46.590
example, you have to know how to do that.

47:47.930 --> 47:54.690
And they have to know, okay, now that they are reconfigured, now how

47:54.690 --> 47:57.670
they have to find a connection to the other agents in the system.

47:58.010 --> 48:01.630
So if you have a set of sensors thrown somewhere and they're

48:01.630 --> 48:06.730
communicating together to achieve some global behavior, they have to

48:06.730 --> 48:10.430
reinstall themselves and then they know how they find a connection,

48:10.550 --> 48:13.130
what they had learned before in the system.

48:15.490 --> 48:17.730
So we have two more systems.

48:17.890 --> 48:21.030
One of them is adaptive systems and the other one is static systems.

48:21.490 --> 48:24.750
You might very well know now what is a static system is.

48:24.770 --> 48:28.250
A static system is that you install it and you have everything in

48:28.250 --> 48:28.470
there.

48:28.710 --> 48:31.490
And it is static, so there is no change in the system.

48:32.130 --> 48:34.490
We don't have any of the cells' X properties.

48:35.010 --> 48:40.190
In adaptive systems, it might have several of these cells' X

48:40.190 --> 48:42.130
properties that we give.

48:42.190 --> 48:45.890
So adaptive systems, you had it also in chapter two or three about

48:45.890 --> 48:47.870
self -organization on adaptive systems.

48:48.430 --> 48:53.830
And you see that an adaptive system can have several properties,

48:54.230 --> 48:55.830
several cells' X properties.

48:56.450 --> 49:00.430
They must have this ability to adapt to the environment.

49:02.390 --> 49:06.570
And there is another word here to self-modify, so how they modify

49:06.570 --> 49:12.050
themselves into different system states.

49:13.410 --> 49:17.990
And there are different ways, like to navigate function in different

49:17.990 --> 49:18.510
environments.

49:18.930 --> 49:22.150
So adaptive system means that your environment is changing.

49:22.790 --> 49:28.130
Like this robot that comes here, and we permanently change this

49:28.130 --> 49:31.630
environment, then it has to know how it adapts to the environment.

49:32.050 --> 49:35.550
So it can have several cells' X properties, and then he has to have

49:35.550 --> 49:38.730
kind of communication with the environment to see if the environment

49:38.730 --> 49:42.510
is changing, how he can change its behavior.

49:43.050 --> 49:47.350
The static system doesn't have any properties that we told you,

49:48.030 --> 49:52.270
doesn't have self-repair or self-detection, protection, and all this

49:52.270 --> 49:52.750
stuff.

49:53.010 --> 49:54.270
There is nothing in there.

49:54.330 --> 49:59.410
There is no self-modification, and they are not able to adapt

49:59.410 --> 50:01.110
environmental scenarios.

50:01.170 --> 50:06.090
We put it here because we wanted to compare it with adaptive systems.

50:06.330 --> 50:11.450
So what is adaptive and a static system in the context of multi-agent

50:11.450 --> 50:11.850
systems?

50:13.410 --> 50:13.930
Good.

50:14.190 --> 50:15.850
So now I finished this chapter.

50:16.610 --> 50:22.950
Thank God, this is the only chapter that we had like more writings and

50:22.950 --> 50:24.210
text and so on.

50:24.670 --> 50:28.870
From now on, it gets, for me, more interesting, I hope, for you.

50:29.190 --> 50:34.090
But what we have here, we have different aspects about cells' X

50:34.090 --> 50:34.590
properties.

50:34.810 --> 50:38.690
We have coordination mechanisms, verification methods, and some cells'

50:38.810 --> 50:39.390
X properties.

50:40.950 --> 50:47.630
And now I want to talk about machine learning and how we can achieve

50:47.630 --> 50:54.310
learning in macroscopic level in organic computing.

50:55.290 --> 50:58.410
Machine learning is an aspect that we already told you in the

50:58.410 --> 50:59.270
architectures.

50:59.530 --> 51:01.690
That is the basic part in organic computing.

51:01.870 --> 51:06.510
We want the system to learn and adapt, for example, to the changing

51:06.510 --> 51:09.310
environment, to anything that is entering to the system.

51:09.370 --> 51:13.470
We want the system to learn, and that's why we want to start now with

51:13.470 --> 51:14.230
the learning.

51:14.970 --> 51:17.190
Do you have any questions about cells' X properties?

51:19.370 --> 51:25.250
So, if I now tell you, give me three names about cells' X, and you can

51:25.250 --> 51:27.830
tell me, self-protection, healing, whatever.

51:28.070 --> 51:32.990
So, in the beginning of this chapter, I asked you, and no one answered

51:32.990 --> 51:33.550
it to me.

51:33.850 --> 51:37.570
But now I think you are able to know what kind of cells' X properties

51:37.570 --> 51:42.130
are there and what are their functionality and so on.

51:51.240 --> 51:55.120
So, now I start with a chapter about machine learning.

51:55.820 --> 52:01.140
We already have the slides on the web page.

52:02.320 --> 52:08.900
We will put the next slides, hopefully this week, online.

52:09.700 --> 52:14.400
There are some parts that we want to change from last year.

52:14.740 --> 52:19.620
That's why you only have the first, I think, 30 slides online, and the

52:19.620 --> 52:20.920
next slides will be online.

52:21.060 --> 52:25.580
Parts 2 and 3 will be online, hopefully, in the next month.

52:26.460 --> 52:33.860
So, who of you has already had a lecture about machine learning?

52:37.270 --> 52:39.430
Okay, what did you have in your course?

52:45.570 --> 52:45.890
Yes.

52:48.490 --> 52:49.910
Neural networks, yes.

52:50.490 --> 52:53.270
Clustering algorithms, mm-hmm.

52:55.210 --> 52:56.490
Data mining.

52:57.810 --> 52:59.110
Okay, good.

52:59.790 --> 53:04.730
So, it could be boring a little bit for you, but whoever plays with a

53:04.730 --> 53:10.450
computer game, with a real computer, not with a computer game, which

53:10.450 --> 53:15.450
could kind of, so it was a game that was kind of learning you.

53:15.950 --> 53:18.090
Like, who played poker online?

53:19.270 --> 53:20.750
Who played chess?

53:22.750 --> 53:26.570
So, not online, I mean computer games.

53:27.530 --> 53:28.250
Chess online?

53:29.550 --> 53:32.670
And you played poker, you said, or chess?

53:34.010 --> 53:38.030
So, these are all, I don't know if that program that you played with

53:38.030 --> 53:43.090
was learning you, but you can write programs that can learn people.

53:43.670 --> 53:50.730
So, that's very, well, I don't know if it makes sense to write

53:50.730 --> 53:56.630
programs who play and learn playing, but that's one of the very, if

53:56.630 --> 54:01.490
you open books in the area of machine learning and read articles in

54:01.490 --> 54:05.950
machine learning, you will see that the very applications are in the

54:05.950 --> 54:08.770
area of games and so on.

54:09.030 --> 54:14.510
So, the first machine learning, the first game was tic-tac-toe.

54:14.710 --> 54:15.970
We will talk about it today.

54:16.410 --> 54:17.410
It's been solved.

54:17.410 --> 54:21.770
So, if you play tic-tac-toe, you play not with a database, but you

54:21.770 --> 54:24.970
play with a real system that has learned that.

54:26.150 --> 54:34.810
The point is, in the 90s, IBM brought a computer game playing chess

54:34.810 --> 54:37.230
and it played against Garry Kasparov.

54:37.370 --> 54:42.490
Garry Kasparov was at that time a Russian chess player and the

54:42.490 --> 54:44.050
computer's name was Deep Blue.

54:44.330 --> 54:49.730
What the computer did was actually nothing than a big database of all

54:49.730 --> 54:53.490
the possible moves, and every time that Garry Kasparov was playing,

54:55.150 --> 54:59.790
the computer was just looking in its database for the supercomputer

54:59.790 --> 55:05.430
built by IBM, and it was like looking what kind of option it could

55:05.430 --> 55:05.830
use.

55:06.530 --> 55:08.230
And this is not a machine learning.

55:08.390 --> 55:10.630
Machine learning is that you learn the system.

55:10.770 --> 55:15.630
Like, if you have a baby, you might have, the baby must have lots of

55:15.630 --> 55:19.030
information about the system, but at some point, it reduces the

55:19.030 --> 55:22.290
information and has the learn mechanism in his brain.

55:22.910 --> 55:27.550
And that's why we want to have a system that learns and not that it's

55:27.550 --> 55:30.770
only a large database looking, searching in the database.

55:32.210 --> 55:38.090
We human beings, we want to learn something in our brain and we just

55:38.090 --> 55:42.970
do it without thinking and having a big database of all possible

55:42.970 --> 55:43.930
options, what we do.

55:43.990 --> 55:45.870
We just learn it and do it.

55:46.430 --> 55:50.130
And that's why we want to see here what kind of options we have in

55:50.130 --> 55:55.110
machine learning, and we talk about the systems, and we kind of, we

55:55.110 --> 55:58.750
come back to the organic computing and say, where in organic computing

55:58.750 --> 56:01.750
do we really need learning and why we are talking all about this

56:01.750 --> 56:03.170
machine learning mechanism.

56:04.190 --> 56:10.930
So, in this chapter, we will have these, hopefully, a few minutes, if

56:10.930 --> 56:12.610
Professor Schmidt doesn't change a lot.

56:13.110 --> 56:18.550
Well, we will definitely talk about learning classifier systems today.

56:19.690 --> 56:25.210
And for learning classifier systems, you need some operators, and

56:25.210 --> 56:30.210
these operators we take from evolutionary algorithms, which we very

56:30.210 --> 56:31.310
briefly talk about.

56:31.390 --> 56:35.690
So, if you are interested more about these kind of algorithms, you can

56:35.690 --> 56:38.730
visit the other course about nature-inspired algorithms.

56:39.330 --> 56:44.950
And then we have neural networks as an alternative to learning

56:44.950 --> 56:46.010
classifier systems.

56:46.650 --> 56:50.510
And we also talk a little bit about swarm intelligence and

56:50.510 --> 56:53.350
particularly ant colony optimization in this chapter.

56:56.890 --> 57:00.110
So, the question is why machine learning in organic computing?

57:00.490 --> 57:04.050
A very quick answer to that is that because we really deal with

57:04.050 --> 57:10.010
different learning aspects and issues, like online and offline

57:10.010 --> 57:11.450
learning for our systems.

57:11.750 --> 57:15.310
So, it's important to know what kind of mechanisms are there and which

57:15.310 --> 57:19.030
one of them we can use in organic computing.

57:21.890 --> 57:24.510
So, there are two aspects that you will learn here.

57:24.630 --> 57:26.450
The first one, definitely, machine learning.

57:26.450 --> 57:32.390
And the other one is that many of these mechanisms you can use in the

57:32.390 --> 57:34.170
area of computational intelligence.

57:34.410 --> 57:36.350
So, it's much more than machine learning.

57:36.690 --> 57:41.430
Like evolutionary algorithms, you can use that for optimizing problems

57:41.430 --> 57:43.150
or solving optimization problems.

57:43.190 --> 57:47.730
It's not only for using in the area of learning.

57:48.750 --> 57:51.590
So, let's start with machine learning.

57:52.350 --> 57:55.250
And we have a machine which wants to learn.

57:57.090 --> 58:06.490
So, this machine, which wants to learn, needs two important things.

58:07.030 --> 58:11.110
And these two important things are the only things that we need to

58:11.110 --> 58:11.450
learn.

58:11.770 --> 58:14.070
The first one is a policy.

58:14.950 --> 58:18.730
So, if you have a policy, what is a policy?

58:19.950 --> 58:23.930
It defines the agent's way of behaving at a given time.

58:25.930 --> 58:27.350
It could be a mapping.

58:27.830 --> 58:28.830
What does that mean?

58:28.930 --> 58:33.470
So, imagine that this is a baby and doesn't know anything in this

58:33.470 --> 58:33.790
world.

58:33.970 --> 58:34.990
And he's hungry.

58:35.810 --> 58:36.750
He cries.

58:37.530 --> 58:40.770
So, situation, hungry, crying.

58:42.810 --> 58:47.250
Then, if he gets food, he has learned it.

58:47.510 --> 58:50.910
Next time he's hungry, he will cry and he will get food.

58:51.690 --> 58:54.970
So, it is as simple as you can think.

58:55.370 --> 59:00.090
So, what happens is that every agent has something, a policy.

59:00.510 --> 59:01.830
And a policy could be a mapping.

59:01.990 --> 59:03.090
Could be something else.

59:03.170 --> 59:04.590
But imagine that this is a mapping.

59:05.230 --> 59:06.810
Situation, behavior.

59:07.190 --> 59:08.410
Situation, action.

59:08.750 --> 59:12.010
So, in any situation that you are, you have an action.

59:12.650 --> 59:16.070
And for this action, you get a reward from the environment.

59:17.030 --> 59:24.090
So, if the baby has kind of, he's hungry, cries and gets food, a

59:24.090 --> 59:27.750
reward function of 100 from the parents, then it learned it.

59:27.890 --> 59:32.630
So, it knows for if hunger happens, if the situation hunger happens,

59:33.270 --> 59:37.510
he cries, reward function from the environment is maximum.

59:38.070 --> 59:42.010
So, a reward function, so it's the second element for learning.

59:42.690 --> 59:44.270
It's a reward function.

59:46.410 --> 59:48.710
And it's some information from the environment.

59:49.330 --> 59:53.790
So, if you play chess and you win, or if you do a move and you win,

59:54.450 --> 59:57.930
you know, oh, that was a good move, you get a high reward function

59:57.930 --> 59:58.690
from the environment.

59:59.370 --> 01:00:00.970
And then you learn it.

01:00:01.070 --> 01:00:02.750
So, it's the next time you do the same thing.

01:00:03.150 --> 01:00:07.930
If you are a baby and you cry and you are hungry and you cry and you

01:00:07.930 --> 01:00:12.150
don't get anything from the environment, then you don't cry next time.

01:00:12.190 --> 01:00:15.210
Maybe you do something else to get food from your parents, right?

01:00:15.310 --> 01:00:19.230
So, it is the same thing happens in the machine learning area.

01:00:19.770 --> 01:00:24.790
So, you need for whatever you want to learn a reward function and your

01:00:24.790 --> 01:00:29.890
goal, your ultimate goal is to maximize this reward function from the

01:00:29.890 --> 01:00:30.310
environment.

01:00:31.970 --> 01:00:35.570
So, here you see we have, well, of course, it's not only status of

01:00:35.570 --> 01:00:37.150
being hungry or whatever.

01:00:37.150 --> 01:00:42.650
If you want to learn a system, you have a long list of situations, a

01:00:42.650 --> 01:00:43.850
long list of actions.

01:00:44.530 --> 01:00:47.590
Each action is related, of course, to one state.

01:00:48.930 --> 01:00:51.950
And for each of them, we have a reward function.

01:00:52.990 --> 01:00:55.030
You want to maximize the reward.

01:00:55.290 --> 01:00:59.430
It means you want to have the maximum reward so that for some of the

01:00:59.430 --> 01:01:04.630
situations, you see that for some state and action, you have no reward

01:01:04.630 --> 01:01:06.090
from the outside.

01:01:06.290 --> 01:01:10.390
So, you know, that wasn't a good kind of state action.

01:01:11.970 --> 01:01:20.730
So, the very first game who was being learned by a very famous

01:01:20.730 --> 01:01:23.750
scientist called David Fogel.

01:01:24.290 --> 01:01:28.210
And David Fogel has written a book called Blondie 24.

01:01:29.450 --> 01:01:33.250
And I suggest everyone who wants to read something about computational

01:01:33.250 --> 01:01:39.590
science and these games, read that book, Blondie 24.

01:01:39.790 --> 01:01:44.690
It's the name of a computer, Blondie, who tries to learn checkers.

01:01:44.850 --> 01:01:46.650
You know checkers, the game checkers?

01:01:46.710 --> 01:01:50.090
It's like chess, but you only have one coin.

01:01:50.630 --> 01:01:54.690
So, it's not that you have different coins, you have simpler than

01:01:54.690 --> 01:01:55.050
chess.

01:01:55.650 --> 01:01:59.210
And David Fogel in his PhD thesis, he works on this tic-tac-toe.

01:01:59.790 --> 01:02:01.130
You know tic-tac-toe?

01:02:02.590 --> 01:02:04.070
Do you play tic-tac-toe?

01:02:05.730 --> 01:02:06.150
Yes?

01:02:06.770 --> 01:02:09.170
Is there anyone who never played tic-tac-toe?

01:02:10.470 --> 01:02:11.550
Okay, so you all know.

01:02:11.810 --> 01:02:13.710
So, now you know what you have to play.

01:02:13.790 --> 01:02:19.150
If you play O, you are lost, right?

01:02:19.930 --> 01:02:20.650
Why?

01:02:20.850 --> 01:02:27.590
Because anywhere you put the O, you can't win.

01:02:27.950 --> 01:02:29.370
So, you are the player here.

01:02:30.050 --> 01:02:31.530
You can't win, right?

01:02:33.690 --> 01:02:37.930
So, because anywhere, so we know it immediately, anywhere it puts...

01:02:37.930 --> 01:02:43.130
So, the cross one has played three times cross, he is playing O.

01:02:43.510 --> 01:02:49.310
Anywhere he puts O, well, anywhere, so he has only one point to put.

01:02:49.410 --> 01:02:53.530
If he puts here, the other one will put here and will win.

01:02:54.970 --> 01:03:01.550
If he puts here, the other one will put here or here, will win.

01:03:02.470 --> 01:03:06.510
Any step he does, he will lose the game.

01:03:07.210 --> 01:03:09.130
But he doesn't know it, so we want to learn it.

01:03:09.130 --> 01:03:14.990
What happens is that he plays, he loses, he gets a reward function

01:03:14.990 --> 01:03:16.670
from the environment that he lost.

01:03:16.810 --> 01:03:18.190
So, reward function was zero.

01:03:19.230 --> 01:03:23.350
So, and then he plays again, he gets another reward function, which is

01:03:23.350 --> 01:03:23.670
zero.

01:03:24.430 --> 01:03:27.690
He puts that in his memory, plays again.

01:03:28.090 --> 01:03:29.610
So, why am I playing again?

01:03:29.710 --> 01:03:31.930
Because the computer doesn't know how to play.

01:03:32.050 --> 01:03:36.810
He plays from the beginning, put it in the first, second, and then the

01:03:36.810 --> 01:03:38.010
third one he is changing.

01:03:39.950 --> 01:03:48.170
And so he wants to have reward 100, he plays again and again, and he

01:03:48.170 --> 01:03:49.450
loses all the time.

01:03:49.690 --> 01:03:51.370
So, what is the conclusion here?

01:03:52.150 --> 01:03:56.170
That was the first two moves, they were wrong.

01:03:57.090 --> 01:04:00.550
Because the third move doesn't matter what kind of third move you do,

01:04:00.630 --> 01:04:01.850
you lose all the time.

01:04:02.190 --> 01:04:07.290
So, the conclusion is that he learned, okay, the first two moves were

01:04:07.290 --> 01:04:11.810
wrong, he has to see how to change the first two moves, that it

01:04:11.810 --> 01:04:14.030
doesn't happen again.

01:04:14.330 --> 01:04:15.910
So, that was wrong.

01:04:16.370 --> 01:04:19.490
Starting from here and putting again here, it's wrong.

01:04:19.970 --> 01:04:22.150
So, at least you have to play like this.

01:04:22.490 --> 01:04:27.350
I mean, we have played all tic-tac-toe, we know if you put two, one

01:04:27.350 --> 01:04:29.370
after the other, we will definitely lose.

01:04:30.490 --> 01:04:34.150
So, the point is that he has to know, okay, he has to change his

01:04:34.150 --> 01:04:34.890
policy.

01:04:35.430 --> 01:04:35.630
Yes?

01:04:42.660 --> 01:04:50.240
Yeah, well, that's just a simple example.

01:04:50.920 --> 01:04:54.480
I just want to say what are the features we need here.

01:04:54.920 --> 01:05:00.120
So, we have now a diploma thesis, writing about poker.

01:05:01.500 --> 01:05:06.460
Well, he finished his diploma thesis, just he sent me the thesis

01:05:06.460 --> 01:05:09.260
yesterday, and it's so complex.

01:05:09.300 --> 01:05:11.160
It's so complex, I mean, you cannot...

01:05:11.160 --> 01:05:12.940
So, poker, you anyway forget it.

01:05:13.460 --> 01:05:17.880
The game's like back and forever you have a chance, then forget it,

01:05:17.980 --> 01:05:21.940
because then there is something, a parameter called chance, that you

01:05:21.940 --> 01:05:22.580
cannot implement.

01:05:23.060 --> 01:05:28.300
But in this such simple example that you just try every field, that's

01:05:28.300 --> 01:05:31.140
actually being solved by David.

01:05:31.260 --> 01:05:34.940
And the game always wins, the computer.

01:05:35.820 --> 01:05:40.260
So, there is no option that the computer never wins against human

01:05:40.260 --> 01:05:40.620
beings.

01:05:41.240 --> 01:05:44.740
But the other chess and so on, they are not...

01:05:44.740 --> 01:05:46.480
Chess, I think it's being solved now.

01:05:47.020 --> 01:05:53.740
It can learn, and checkers as well, but complex games not.

01:05:53.860 --> 01:05:56.840
And complex environments, forget it anyway.

01:05:59.640 --> 01:06:00.600
So, conclusion.

01:06:01.800 --> 01:06:04.400
First of all, the agent starts with no knowledge.

01:06:04.800 --> 01:06:11.240
So, the more you let it run, so you try 10,000 times, or in this case

01:06:11.240 --> 01:06:15.080
20 times or 30 times, then it has more experience.

01:06:15.440 --> 01:06:17.780
And with more experience, he can learn more.

01:06:18.320 --> 01:06:19.600
What we need is that

01:06:22.860 --> 01:06:25.740
we want to learn to obtain a good reward.

01:06:25.880 --> 01:06:29.180
We don't want to learn how to put it, so we want to learn how we

01:06:29.180 --> 01:06:30.360
obtain a good reward.

01:06:30.860 --> 01:06:36.100
You don't want to have a large database and large experience memory to

01:06:36.100 --> 01:06:43.120
search every time how we put a sign in tic-tac-toe, but we want to see

01:06:43.120 --> 01:06:45.220
how we can achieve a high reward.

01:06:45.520 --> 01:06:50.980
So, that must be saved, not a large database about all different

01:06:50.980 --> 01:06:51.400
moves.

01:06:52.600 --> 01:06:54.920
We use experience for that.

01:06:55.920 --> 01:07:01.220
We can also predict in this way what could be the next move of the

01:07:01.220 --> 01:07:02.960
player that you can also do.

01:07:04.160 --> 01:07:08.080
And you can also estimate the reward, so you estimate the reward

01:07:08.080 --> 01:07:09.200
before you play it.

01:07:11.780 --> 01:07:14.500
And, of course, the experience you can always update.

01:07:14.680 --> 01:07:18.860
So, these are the conclusions, and you will see many of these will

01:07:18.860 --> 01:07:22.000
appear in machine learning in the next slides.

01:07:23.140 --> 01:07:26.400
So, now let's start with... oh, okay.

01:07:26.520 --> 01:07:27.460
So, now we have here

01:07:30.560 --> 01:07:33.000
categories of different machine learning methods.

01:07:33.980 --> 01:07:36.460
The first one is about reinforcement learning.

01:07:37.520 --> 01:07:42.680
Reinforcement learning you can use to learn exactly the same problem I

01:07:42.680 --> 01:07:44.980
showed you by temporal difference learning.

01:07:45.060 --> 01:07:46.780
There you have prediction mechanisms.

01:07:47.040 --> 01:07:48.980
You look in the future.

01:07:49.200 --> 01:07:54.640
So, here with temporal, with TDL, you have prediction mechanisms.

01:07:55.000 --> 01:07:58.480
With Bayesian learning, you have a large database, and you always

01:07:58.480 --> 01:07:59.740
compute the probabilities.

01:08:00.480 --> 01:08:07.520
So, you kind of put your knowledge in a graph, and then for each move

01:08:07.520 --> 01:08:11.160
that is followed in your graph, you compute a probability, and you

01:08:11.160 --> 01:08:14.380
look for the node in your system that has the highest probability.

01:08:14.860 --> 01:08:16.120
We don't talk about them here.

01:08:16.220 --> 01:08:20.280
That's why we just put the names that you know that these are the

01:08:20.280 --> 01:08:21.320
methods that exist.

01:08:26.540 --> 01:08:27.200
Probabilities.

01:08:29.200 --> 01:08:34.420
So, we now talk about learning classifier systems and neural networks.

01:08:35.520 --> 01:08:42.220
And for them, working with learning classifier systems, we will take

01:08:42.220 --> 01:08:45.520
some aspects from evolutionary algorithm and swarm intelligence.

01:08:46.260 --> 01:08:48.540
Swarm intelligence is a large area.

01:08:48.640 --> 01:08:53.200
We have two topics in that, two major topics in that area, and we talk

01:08:53.200 --> 01:08:57.440
about ant colony optimization, because you already had that in your

01:08:57.440 --> 01:09:02.200
other lectures, so that would be easier for you to work on that.

01:09:02.980 --> 01:09:08.080
So, now I want to work with learning classifier systems.

01:09:08.680 --> 01:09:14.160
And for that, I want to change the example that I had and talk about

01:09:14.160 --> 01:09:18.320
another simpler, but in a way more complex example.

01:09:19.760 --> 01:09:30.620
Has any one of you had this multiplexer in his studies, or her

01:09:30.620 --> 01:09:31.560
multiplexer?

01:09:33.960 --> 01:09:38.140
So, this is an example about multiplexer, but it's very simple.

01:09:38.280 --> 01:09:43.580
So, a multiplexer is nothing than a bit string.

01:09:43.900 --> 01:09:46.740
So, you have a bit string of length k.

01:09:47.580 --> 01:09:51.140
So, a bit string of length k could be something like this.

01:09:54.810 --> 01:09:59.970
But why we call it multiplexer is because this has two parts, this bit

01:09:59.970 --> 01:10:00.350
string.

01:10:00.750 --> 01:10:02.930
The first part and the second part.

01:10:03.290 --> 01:10:08.090
The first part, mainly this one, this two,

01:10:12.120 --> 01:10:14.860
it has a length of l.

01:10:15.260 --> 01:10:20.540
So, here l is equal to two, two elements.

01:10:21.300 --> 01:10:27.160
And the second part has a length which depends on the l.

01:10:27.580 --> 01:10:35.780
So, this is, the second part has l, two to the power of l, two to the

01:10:35.780 --> 01:10:38.060
power of two, four elements.

01:10:39.460 --> 01:10:43.980
So, any bit string I tell you, I give you, must have, if it has this

01:10:43.980 --> 01:10:47.760
property, you can say that it's like a multiplexer.

01:10:49.040 --> 01:10:53.600
So, it's a bit string of certain length.

01:10:54.620 --> 01:10:58.980
And the first part, so if you measure the length, and the first part

01:10:58.980 --> 01:11:04.920
has, let's say, l elements, the second part must have two to the power

01:11:04.920 --> 01:11:05.380
of l.

01:11:05.700 --> 01:11:12.200
Then this bit string, you can call it multiplexer, if it has, of

01:11:12.200 --> 01:11:14.080
course, other properties that I will tell you.

01:11:14.940 --> 01:11:19.920
But you get it here, like, you can have another, so you can have here

01:11:19.920 --> 01:11:22.640
three elements, or here zero, one, one, one.

01:11:23.140 --> 01:11:25.560
And the next part must have how much?

01:11:26.460 --> 01:11:29.700
Two to the power of three, eight elements.

01:11:32.520 --> 01:11:33.040
Right?

01:11:33.220 --> 01:11:39.260
If he has here three, he must have two to the power of three elements,

01:11:39.460 --> 01:11:41.440
so it will be one, zero, zero, whatever.

01:11:46.480 --> 01:11:48.980
This will be a new bit string.

01:11:52.180 --> 01:11:53.120
You get it?

01:11:54.260 --> 01:11:59.720
Multiplexer, so you have a bit string, and then you measure the length

01:11:59.720 --> 01:12:02.980
of this bit string, which will be k, okay?

01:12:03.280 --> 01:12:08.560
And in the beginning you have l elements, and then the property of

01:12:08.560 --> 01:12:12.880
this bit string is that in the beginning you have l, in the second one

01:12:12.880 --> 01:12:15.060
you have two to the power of l.

01:12:16.500 --> 01:12:23.580
So, if I tell you k-bit multiplexer, then you know, okay, you have k

01:12:23.580 --> 01:12:27.140
bits, and it must have actually some certain sizes.

01:12:27.260 --> 01:12:29.480
You cannot have five-bit multiplexer.

01:12:30.360 --> 01:12:34.220
You must have certain size to have the multiplexer for a bit string,

01:12:34.360 --> 01:12:38.060
because the first part and second part, they have a relationship with

01:12:38.060 --> 01:12:38.700
each other.

01:12:39.960 --> 01:12:42.020
So, and now what multiplexer is?

01:12:42.920 --> 01:12:47.340
Multiplexer is that you give this bit string, the entire complete bit

01:12:47.340 --> 01:12:52.160
string, into this multiplexer, and a value comes out.

01:12:53.680 --> 01:12:54.700
And what is this value?

01:12:55.140 --> 01:12:57.860
This value is defined in this bit string.

01:12:59.120 --> 01:13:07.140
This value says you give... so the first part is actually an address.

01:13:08.240 --> 01:13:12.420
So, one zero, a bit, one zero means two.

01:13:13.340 --> 01:13:14.460
One zero is two.

01:13:14.680 --> 01:13:15.780
Zero zero is zero.

01:13:16.020 --> 01:13:17.100
Zero one is one.

01:13:17.720 --> 01:13:18.900
One zero is two.

01:13:19.080 --> 01:13:20.000
One one is three.

01:13:20.240 --> 01:13:21.280
It's a bit string, right?

01:13:21.920 --> 01:13:27.180
So, this says two, and the other one, you go here, it has addresses.

01:13:27.340 --> 01:13:30.240
So, address zero, one, two, and three.

01:13:31.100 --> 01:13:37.840
You take the bit from that address, which is shown in the first part.

01:13:38.180 --> 01:13:39.400
So, this is address two.

01:13:39.840 --> 01:13:45.440
You come here to the first part, select the element from address two,

01:13:45.860 --> 01:13:49.040
and that will be the output of the multiplexer.

01:13:52.120 --> 01:14:00.160
So, if I have now address... if I have a bit of like zero, zero, one,

01:14:00.320 --> 01:14:03.160
zero, one, one, what will be the output?

01:14:05.020 --> 01:14:05.580
One.

01:14:09.120 --> 01:14:16.520
If I have one, one, one, zero, one, one, one,

01:14:20.270 --> 01:14:21.670
say something, one.

01:14:22.090 --> 01:14:22.810
Thank you.

01:14:24.850 --> 01:14:25.370
Good.

01:14:25.470 --> 01:14:32.220
If I have... what will be the value?

01:14:33.200 --> 01:14:33.840
Zero.

01:14:34.260 --> 01:14:34.620
Why?

01:14:35.120 --> 01:14:39.600
Because this is showing the element zero, the address zero, and

01:14:39.600 --> 01:14:41.320
address zero is here, zero.

01:14:46.620 --> 01:14:49.040
So, I have to remove this here.

01:14:49.160 --> 01:14:50.900
I think this is also in the...

01:14:53.460 --> 01:14:55.060
So, this is a multiplexer.

01:14:55.160 --> 01:14:58.940
So, you put this bit string, which has a certain property into a

01:14:58.940 --> 01:15:05.740
multiplexer, and output, a bit comes out, which is the other end.

01:15:05.860 --> 01:15:08.280
You can use... you can see that you can use that in the computer.

01:15:08.840 --> 01:15:12.700
You give a certain address of the memory, and you get the input of

01:15:12.700 --> 01:15:13.480
that memory out.

01:15:13.740 --> 01:15:16.080
That's why you need the multiplexer in computer systems.

01:15:17.140 --> 01:15:22.620
Look at your Infosy documents, and you will find out that there you

01:15:22.620 --> 01:15:24.340
had actually multiplexer.

01:15:27.500 --> 01:15:27.980
Okay.

01:15:27.980 --> 01:15:30.160
So, this I can define it as a rule.

01:15:30.320 --> 01:15:37.540
I can say whatever I have in the beginning, input, multiplexer, double

01:15:37.540 --> 01:15:39.540
point, output of the multiplexer.

01:15:40.320 --> 01:15:40.700
Right?

01:15:40.900 --> 01:15:42.500
So, a rule, I define it.

01:15:42.500 --> 01:15:49.180
It's not actually a complete rule, but a rule, let's say, will be this

01:15:49.180 --> 01:15:49.860
thing.

01:15:50.860 --> 01:15:56.920
So, a bit string, and I call zero the action of multiplexer.

01:16:00.560 --> 01:16:04.920
Now, what we want to do with all this machine learning, we want to

01:16:04.920 --> 01:16:08.220
learn this Boolean function.

01:16:08.380 --> 01:16:09.260
So, this is a function.

01:16:09.380 --> 01:16:10.740
This multiplexer is a function.

01:16:11.340 --> 01:16:12.660
We want to learn it.

01:16:12.980 --> 01:16:14.520
And now, why we want to learn it?

01:16:14.580 --> 01:16:17.800
So, what is the difficult thing in that we want to learn it?

01:16:19.100 --> 01:16:23.460
Because if we have... so, that I showed you a very small string.

01:16:23.880 --> 01:16:27.360
Usually, we have strings of very large size.

01:16:28.200 --> 01:16:32.280
And if you want to find what is the output, then you have lots of

01:16:32.280 --> 01:16:35.380
options to go through to see what is the output of that system.

01:16:36.060 --> 01:16:40.620
So, imagine if k plus one, k was the length, and k plus one we

01:16:40.620 --> 01:16:44.580
consider because we also consider the action to what is that length.

01:16:45.460 --> 01:16:48.480
If you have four, we have 16 options to go through.

01:16:48.620 --> 01:16:49.300
That's easy.

01:16:50.280 --> 01:16:55.380
If you have 30, that is a normal, not even a normal size.

01:16:55.520 --> 01:16:56.660
It's a very small size.

01:16:57.100 --> 01:17:01.980
You have so many options to check if the multiplexer is correct or

01:17:01.980 --> 01:17:02.320
wrong.

01:17:02.880 --> 01:17:06.840
So, that's a good reason to learn this system, because it's getting

01:17:06.840 --> 01:17:07.420
very complex.

01:17:07.580 --> 01:17:11.940
So, if you have a length of hundred, then you need a lot of time to

01:17:11.940 --> 01:17:17.020
check all this to learn it, to check everything and do an exhaustive

01:17:17.020 --> 01:17:17.340
search.

01:17:17.540 --> 01:17:20.040
So, it's better to learn it.

01:17:21.680 --> 01:17:23.520
So, what can we do now?

01:17:26.180 --> 01:17:28.360
What we do is that we are a computer.

01:17:29.400 --> 01:17:30.840
We want to learn multiplexer.

01:17:30.940 --> 01:17:35.600
We want to write a program which gives us a perfect multiplexer by

01:17:35.600 --> 01:17:37.860
learning this rule that we gave.

01:17:38.540 --> 01:17:41.980
So, what we define, we define what do we need is a reward function.

01:17:42.280 --> 01:17:48.300
So, every time I have an input, I have an action, I see if that was

01:17:48.300 --> 01:17:49.160
the correct one or not.

01:17:49.220 --> 01:17:49.980
I can always correct.

01:17:50.080 --> 01:17:53.140
I can check the address and see if it was the correct one or not.

01:17:53.340 --> 01:17:57.280
And I can get from the environment a reward function, which the

01:17:57.280 --> 01:18:01.020
highest reward function will be, let's say, a hundred, and the lowest

01:18:01.020 --> 01:18:01.800
one will be zero.

01:18:01.980 --> 01:18:04.860
So, if it is wrong, we'll get a very bad reward function.

01:18:05.020 --> 01:18:08.240
If it is correct, it will get something like hundred.

01:18:08.860 --> 01:18:13.980
So, here in this case, you see here, this is address two.

01:18:14.420 --> 01:18:15.700
We are looking for zero.

01:18:16.640 --> 01:18:18.020
It was correct.

01:18:18.200 --> 01:18:18.820
We get hundred.

01:18:19.300 --> 01:18:20.200
Here it is wrong.

01:18:20.280 --> 01:18:21.600
We get reward of zero.

01:18:25.680 --> 01:18:33.580
So, now we define input, action, and reward as a rule.

01:18:33.900 --> 01:18:39.320
So, now I define the rule a little bit bigger with the information

01:18:39.320 --> 01:18:41.180
about the reward function.

01:18:41.480 --> 01:18:43.260
We have now the rule.

01:18:45.240 --> 01:18:46.360
This is a good rule.

01:18:46.720 --> 01:18:50.360
This is a terrible rule because it has a reward of zero.

01:18:54.140 --> 01:19:00.420
In some books and some literature, you can call this reward as a

01:19:00.420 --> 01:19:00.720
payoff.

01:19:01.160 --> 01:19:05.140
So, it's a payoff of this, but we use the reward function all over

01:19:05.140 --> 01:19:11.060
this lecture and just for your information.

01:19:11.860 --> 01:19:17.940
The other thing is that if you have here an address of two, so one,

01:19:18.060 --> 01:19:24.120
zero, what is interesting for us is the one in the address two, which

01:19:24.120 --> 01:19:24.560
is here.

01:19:25.060 --> 01:19:28.820
The rest is actually could be anything else, right?

01:19:29.060 --> 01:19:32.700
That could be zero or one or could be undefined.

01:19:33.580 --> 01:19:37.460
That's why we define something which is called wildcard.

01:19:37.840 --> 01:19:40.000
And wildcard means zero or one.

01:19:40.240 --> 01:19:43.540
So, I can have a rule which looks like this,

01:19:53.340 --> 01:19:59.520
and still I can give a reward of 100 because I don't care what are the

01:19:59.520 --> 01:20:03.700
other elements that is correct, and then I can give it a reward

01:20:03.700 --> 01:20:04.580
function of 100.

01:20:05.380 --> 01:20:09.840
And a wildcard could be zero or one.

01:20:10.000 --> 01:20:12.100
So, you say this is undefined.

01:20:12.200 --> 01:20:15.340
And this is very useful if you have large... yes?

01:20:33.330 --> 01:20:37.890
Yes, but you know, it's better to... later you will see that we will

01:20:37.890 --> 01:20:42.550
have large bit strings and some of them are unknown because they are

01:20:42.550 --> 01:20:43.630
not being learned yet.

01:20:43.890 --> 01:20:48.590
But still, if we learn it, then we will put some correct variable

01:20:48.590 --> 01:20:48.910
there.

01:20:49.210 --> 01:20:54.130
But in some cases, you can have exactly two rules with wildcards, but

01:20:54.130 --> 01:20:56.710
then still you can use the information there.

01:20:56.950 --> 01:21:03.230
If you stay a little bit with me, then I explain it later how that is

01:21:03.230 --> 01:21:03.690
wildcard.

01:21:03.870 --> 01:21:05.930
But now you know the meaning of a wildcard.

01:21:06.030 --> 01:21:07.610
So, wildcard is either zero or one.

01:21:07.690 --> 01:21:08.250
I don't care.

01:21:08.870 --> 01:21:11.190
In electrical engineering, we call it don't care.

01:21:12.490 --> 01:21:12.770
Okay?

01:21:13.610 --> 01:21:15.450
Here, we call it wildcard.

01:21:26.450 --> 01:21:28.490
Yeah, you know, it's...

01:21:29.850 --> 01:21:30.250
Yeah.

01:21:30.510 --> 01:21:33.290
So, you know, this is for modeling, sir.

01:21:33.430 --> 01:21:37.570
So, we are not still learning process.

01:21:37.650 --> 01:21:38.330
We are just modeling.

01:21:38.470 --> 01:21:39.490
So, this is a bit string.

01:21:39.830 --> 01:21:42.010
Some elements we say they are unknown.

01:21:43.070 --> 01:21:47.090
But that doesn't... You can have something like this here.

01:21:54.640 --> 01:21:58.260
So, here could be one or could be zero.

01:21:58.980 --> 01:21:59.760
You don't know.

01:22:00.540 --> 01:22:05.300
So, that rule will be... So, we are modeling now the problem, you

01:22:05.300 --> 01:22:05.440
know?

01:22:05.540 --> 01:22:13.120
And that will be maybe a reward zero because it's unknown.

01:22:14.000 --> 01:22:17.940
So, but here in this way, I just wanted to define a bit that is

01:22:17.940 --> 01:22:18.720
unknown to us.

01:22:18.740 --> 01:22:25.100
And you will see in the next slide how I can use this unknown bit in

01:22:25.100 --> 01:22:27.180
the learning classifier system.

01:22:28.700 --> 01:22:31.660
So, now let's see if we can learn it.

01:22:31.720 --> 01:22:36.820
So, we have this robot here and he wants to learn multiplexer.

01:22:38.100 --> 01:22:43.860
What we assume now is that this robot has already a list of different

01:22:43.860 --> 01:22:44.280
actions.

01:22:44.500 --> 01:22:50.680
How we make this list, it will be a long part about evolutionary

01:22:50.680 --> 01:22:52.220
algorithms to make the list.

01:22:52.400 --> 01:22:59.140
But imagine that this robot has a population of different rules,

01:22:59.760 --> 01:23:00.060
right?

01:23:00.720 --> 01:23:05.340
So, this robot has this population and he already has this in his

01:23:05.340 --> 01:23:08.320
brain somehow in the other experience.

01:23:09.040 --> 01:23:12.480
And here you see you have some of them have wildcards.

01:23:13.120 --> 01:23:18.840
And here some rules have some kind of reward from previous moves that

01:23:18.840 --> 01:23:19.420
he had done.

01:23:20.640 --> 01:23:22.000
So, what happens?

01:23:22.400 --> 01:23:23.560
He has this population.

01:23:24.120 --> 01:23:25.680
Now the input comes in.

01:23:26.980 --> 01:23:29.240
This input comes in.

01:23:29.240 --> 01:23:30.780
He has this population.

01:23:31.340 --> 01:23:32.200
What does he do?

01:23:32.320 --> 01:23:37.620
He compares the input with his knowledge, which is now a population.

01:23:39.060 --> 01:23:44.000
Now, if you look more into details, we have here input is 1, 1, 1, 0,

01:23:44.140 --> 01:23:45.740
1, 1, 1, 0.

01:23:46.900 --> 01:23:51.880
If you compare, you will find that there are two rules, namely this

01:23:51.880 --> 01:23:58.640
one and this one, that are exactly matching the input.

01:24:00.500 --> 01:24:01.140
Right?

01:24:01.400 --> 01:24:10.140
So, 1, 0, 1, 1, 0 will match here and the other one will match here.

01:24:15.460 --> 01:24:19.940
So, these are two rules with two different actions.

01:24:19.940 --> 01:24:25.480
One of them has action 1, the other one has action 0, the different

01:24:25.480 --> 01:24:26.400
reward values.

01:24:27.840 --> 01:24:33.560
So, these two will be taken out of the population, will put in a small

01:24:33.560 --> 01:24:36.420
population, which we call it a match set, because they were matched.

01:24:36.720 --> 01:24:38.020
The others were not matched.

01:24:38.480 --> 01:24:43.960
So, we take those who are matched, we put them in a match set.

01:24:44.820 --> 01:24:47.280
So, now what to do?

01:24:48.400 --> 01:24:50.500
So, now you will say it's easy.

01:24:50.860 --> 01:24:57.540
The agent must see which one of them has the highest reward and then

01:24:57.540 --> 01:24:58.840
he will take that as an action.

01:24:59.300 --> 01:24:59.780
Right?

01:24:59.840 --> 01:25:01.300
So, that will be the easiest thing.

01:25:01.380 --> 01:25:05.060
So, he has an input, he has his knowledge and experience in a

01:25:05.060 --> 01:25:05.600
population.

01:25:06.080 --> 01:25:10.960
What happens is that he sees which one of them has the highest reward,

01:25:12.960 --> 01:25:21.360
namely this one, and he will take this as the action 1, whatever he

01:25:21.360 --> 01:25:21.580
has.

01:25:21.920 --> 01:25:27.040
He applies it to the environment, gets a new reward function back, and

01:25:27.040 --> 01:25:28.500
then he can update it and so on.

01:25:30.120 --> 01:25:34.180
So, well, everything I said before is now here.

01:25:35.700 --> 01:25:38.950
But what happens if he gets...

01:25:38.950 --> 01:25:41.890
So, that's the easy way we have here.

01:25:43.510 --> 01:25:46.690
You have one with a higher fitness value, the other one with a lower

01:25:46.690 --> 01:25:47.310
fitness value.

01:25:47.390 --> 01:25:51.010
So, you take the reward value and then you take the one with the

01:25:51.010 --> 01:25:52.390
highest reward function.

01:25:52.850 --> 01:25:59.670
What happens if you have many rules in your population?

01:25:59.670 --> 01:26:01.130
Now, we change the population.

01:26:01.310 --> 01:26:05.670
Now, imagine that you have so many match sets have changed and you

01:26:05.670 --> 01:26:07.490
have so many rules in your match set.

01:26:07.610 --> 01:26:09.030
So, now imagine it's a new thing.

01:26:09.250 --> 01:26:14.470
It doesn't have to do with the population before, but it's now imagine

01:26:14.470 --> 01:26:17.070
that we have other rules in the match set.

01:26:19.670 --> 01:26:20.870
And now, what should he do?

01:26:21.450 --> 01:26:30.270
So, he has two of them with action 1, both of them with high reward

01:26:30.270 --> 01:26:38.810
values, 67 and 70, and two rules with action 0, one of them with 65,

01:26:39.070 --> 01:26:40.130
the other one with 3.

01:26:40.950 --> 01:26:44.250
What he does actually, he makes an average.

01:26:44.390 --> 01:26:49.690
So, he says if he computes a probability for 0, for taking the action

01:26:49.690 --> 01:26:58.510
0, for taking action A, he makes an average of the reward value, 3 and

01:26:58.510 --> 01:27:08.990
65 for 0, 3 and 65 for 0, over all the reward value for all the

01:27:08.990 --> 01:27:09.330
actions.

01:27:09.970 --> 01:27:11.990
And then, for 1, he does it the same.

01:27:12.990 --> 01:27:15.770
And of course, for 1, this value will be higher.

01:27:16.050 --> 01:27:24.570
So, it will select action equals 1, and that is what he will take.

01:27:27.740 --> 01:27:29.360
So, did you get that part?

01:27:30.100 --> 01:27:31.200
So, very good.

01:27:32.540 --> 01:27:36.440
So, what happens is that he applies this rule to the environment,

01:27:36.780 --> 01:27:40.420
plays it again or whatever, so he applies this multiplexer, he gets a

01:27:40.420 --> 01:27:41.040
reward back.

01:27:41.260 --> 01:27:45.260
And now, he has to update the reward function in his population, which

01:27:45.260 --> 01:27:49.040
was a very small population, but still he has to update the

01:27:49.040 --> 01:27:50.700
information on his list.

01:27:51.180 --> 01:27:51.420
So,

01:27:54.720 --> 01:27:56.520
yeah, and he does it.

01:27:56.560 --> 01:28:00.960
So, you will see how he will do it, but he applies that rule to the

01:28:00.960 --> 01:28:05.220
multiplexer, gets a reward function, and then he will change the

01:28:05.220 --> 01:28:07.420
reward value of that rule in his population.

01:28:08.580 --> 01:28:12.340
The point is that how do we get the population?

01:28:12.500 --> 01:28:14.600
The first population is the important thing.

01:28:14.980 --> 01:28:18.960
How do we give this population to the agent?

01:28:19.400 --> 01:28:22.980
That's the thing that we have to learn, and we cannot keep all the

01:28:22.980 --> 01:28:24.380
rules in his memory.

01:28:24.580 --> 01:28:27.000
We don't have so many space in there.

01:28:27.220 --> 01:28:31.180
We have to see how we make this population, and this is learning in

01:28:31.180 --> 01:28:31.400
there.

01:28:32.140 --> 01:28:37.540
To know which population to keep is the learning aspect that we will

01:28:37.540 --> 01:28:37.840
learn.

01:28:37.840 --> 01:28:42.880
So, what you learn now was the basics of learning classifier systems,

01:28:44.540 --> 01:28:48.200
and oh, we don't have time.

01:28:48.400 --> 01:28:52.900
So, next week, you will learn different aspects of learning classifier

01:28:52.900 --> 01:28:56.520
systems, how we make rules, how we make rewards, and so on.

01:28:56.960 --> 01:29:04.780
And Professor Schmeck will be here, hopefully, and I might be away for

01:29:04.780 --> 01:29:05.480
a conference.

01:29:06.960 --> 01:29:10.360
So, see you then in two weeks, but Professor Schmeck will be here next

01:29:10.360 --> 01:29:10.640
week.

01:29:10.860 --> 01:29:11.400
Thank you.

