For the complete documentation index, see /llms.txt. The full corpus is available at /llms-full.txt. Most pages have a raw markdown variant; append .md to the URL or send Accept: text/markdown.

Invert University  |  Research

Systems Engineering for Biotechnology: Modeling, Optimization, and Control

Dr. Sebastián Espinel Ríos · University College Dublin

Aug 2026 · 52:38 · 6,564-word transcript

About this talk

Dr. Sebastián Espinel Ríos (University College Dublin) introduces the Systems Engineering for Biotechnology group and the problem it is built around. Bioprocess optimization is still largely recipe-driven and open loop, so a golden batch gets locked in and cannot correct itself mid-run. He walks through the methods the group develops in response, including constraint-based and surrogate metabolic models, hybrid models that pair mechanistic knowledge with machine learning, soft sensors that reconstruct intracellular states where no hardware sensor exists, and reinforcement learning for decision policies under uncertainty. For process development, MSAT, and data science teams working on digital twins and advanced process control.

Transcript

6,564 words · 52 min

Automatically transcribed and corrected for technical terms. Timestamps refer to Systems Engineering for Biotechnology: Modeling, Optimization, and Control above.

Thank you very much for the invitation to introduce the group I'm leading at University College Dublin in the School of Chemical and Bioprocess Engineering. And the research group I'm leading there is the Systems Engineering for Biotechnology research group. And before I get started into more technical details of what we do in the group, I would like first to introduce myself with a little bit of my background. So I am originally from Colombia. I was raised in Mexico, and there in Mexico, I did a bachelor's in biotechnology engineering at the Autonomous University of Yucatán. Then I moved to the Netherlands, where I completed a master's degree in biotechnology with a specialization or focus in bioprocess engineering, process technology aspects.

Later on, I moved to Germany, where I joined the Max Planck Institute for Dynamics of Complex Technical Systems in Magdeburg, where alongside, while I was still affiliated with the Otto-von-Guericke University Magdeburg, I completed my PhD in process and systems engineering with a focus on bioprocesses, specifically a concept that I will introduce later called metabolic cybergenetics to make bioprocesses more efficient. Then I continued with two postdoctoral positions, one at Princeton University in the US in the Department of Chemical and Biological Engineering in the Avalos Lab. And there I continued working on digitalization aspects of bioprocess engineering, specifically towards enabling metabolic cybergenetic ideas, as I will introduce later.

I also wanted to experience research towards industry applications, and then I moved to CSIRO, which is Australia's national science agency, where I was part of the Advanced Future Science platform, where we wanted to develop new biotechnological ideas to enable the bioeconomy in Australia. And finally, I landed in Ireland, where, as I said, I'm currently leading the Systems Engineering for Biotechnology research group, and I'm an assistant professor of bioproduction there. I'm also part of the Ad Astra program, so I'm an Ad Astra fellow, which are like new faculty members that UCD consider as a priority to advance its research strategic goals. So with this, let me introduce what we actually do in the group.

So we are focusing on developing tailored computational methods for advanced modeling, optimization, and control towards advancing the digitalization of the biotechnology industry and enabling novel biotechnological applications. And this is driven by three aspects. So one is knowledge abstraction, and let me activate my laser. So one thing is how can we abstract mechanistic metabolic knowledge, but also the bioreactor physics alongside process data that we gather during the operation of these bioprocesses, in order to make sense of our bioprocess operation and use this information to optimize it and control it in a smart way. Naturally, this also involves capturing plant-wide context because we know bioprocesses are not isolated things. They are actually involved within a bigger scope involving upstream and downstream unit operations.

The second thing, as I said, is exploit this knowledge, which I already mentioned a bit. So we want to use this abstracted knowledge and enable advanced control strategies. For example, taking smart decisions in terms of metabolic actuation, feeding strategies as in fed-batch systems, but also optimizing initial conditions or even process design parameters. And ideally, we want to do this in real time in a closed loop in a smart fashion, such that we can take readings of the system, we can know the state of our system to update what we know about it in real time, and therefore update our control actions. And the other thing that we are driven by is knowledge facilitation.

So we are aware that many of the methods that we work and I will introduce also later are actually quite complex for non-expert practitioners in the biotechnology industry. So we want to develop tools that facilitate their adoption by the wider biotechnological community. Naturally, when I mention this knowledge abstraction concept, we are typically talking about what people in the field call digital twins, right? And to enable the abstraction of this knowledge towards its exploitation, we make use of a lot of scientific machine learning tools from hybrid modeling concepts such as neural networks, Gaussian processes, neural ODEs, and so on. What I would like to introduce now is the context of our research.

First of all, biotechnology, and we know that is key for the development of the bio-based circular economy and also for improved human health through the production of a wide variety of value-added products. From a high-value, low-volume product such as pharmaceuticals, cosmetics, to low-value, high-volume products such as gas and fuels. And we have, of course, in between commodities, specialties such as nutrients, chemicals, fertilizers, and materials. So there is a wide palette of products that we can produce through biotechnology, specifically exploiting the metabolism of cell factories, which can be of microbial or mammalian nature. Right? The concept is we have this metabolic network here that determines the potential of my microbial cell factory, and that is often what we will be targeting as the focus of our research and also the novelties in our technological developments.

Another thing to understand before we jump into technical aspects is different sectors in biotechnology have distinct different challenges and also priorities. For example, when we focus on commodity chemicals and liquid and gas fuels, people typically focus on enhancing, or they are, let's say, prioritizing high volumetric productivity. Why? Because in these cases, they are competing directly with petrochemical technologies and resources. So they need to be economically competitive and be equally, let's say, efficient as the petrochemical industry. So in these cases, of course, their priority, as I said, is to maximize this volumetric productivity but also use a substrate, the carbon source, the feedstock efficiently. They don't want to waste it because that's money that they would be wasting and overall be very, very cost-effective.

Now, when we talk about pharmaceuticals and cosmetics, these high-value, low-volume products, profitability, as a thing, is actually driven by critical quality attributes. So it's not really too much about producing a lot, but it's about that what I'm producing has high quality. So I'm talking about vaccines, about recombinant proteins, et cetera. And another thing is naturally the time to market. So we want to be fast from the point where we have a new development and we can go to market. So concepts such as knowledge transfer here becomes really, really important. So in summary here, the priority is, as I said, rapid knowledge transfer and scale up to go to market.

And it has, of course, a priority on critical quality attributes rather than maximizing volumetric productivity. And why is that? Well, here we are not competing directly with petrochemical resources. We cannot produce a vaccine with oil. Right? So as you can see, different sectors will have different challenges and priority. And I also like to say, thinking a little bit forward, that biotechnology expands beyond Earth. Right? So nowadays with human space exploration, companies, national agencies are now working a lot on how to exploit biotechnological technologies towards producing food, medicines, and materials in situ, in space, such that we don't have the burden of moving materials from Earth to space, facilitating or enabling a longer space exploration and colonies as well.

So even though I have given you a very promising picture of biotechnology, I also have to be honest, and that is biotechnology has major challenges still to enable its wider adoption. So mainly it has to do with the fact that these biotechnological processes are still difficult to predict, to optimize, and operate robustly. Right? And there are several reasons that explain this that I will talk next. So what are these main challenges? So one thing that I have identified is, well, there's limited integration of degrees of freedom across scale. What do I mean? So if we talk with process systems engineers, often like chemical and process engineers, when they want to optimize bioprocesses, they would largely focus on the extracellular or macroscopic level, such as optimizing, for instance, the medium formulation, the culture environment, pH, temperature, agitation, the mode of operation, for example, batch, fed-batch, continuous, and bioreactor design parameters such as the impeller and so on.

So if you think about it, we are still talking about things around the cell, not in the cell. So these are focusing on optimizing overall conditions. So in a way, it's like the overall health of my system. However, as I said in the introduction, so the machinery that enables bioprocesses are all these metabolic pathways, the metabolic reaction network that converts substrates into products and byproducts and also enables growth through producing more biomass, more cells. So another, let's say, school would be systems biologists. So these are people who actually try to target metabolism directly through metabolic engineering and genetic engineering aspects. So in here, we need to understand that wild-type microorganisms have been, let's say, engineered by evolution, if you like, to maximize biomass production. Why?

Because, well, this makes sense from an evolutionary point of view. You want to produce as many new cells as you can. This typically means that products, kind of waste materials from the point of view of the cell, are produced in very little quantities. And the funny thing is actually these products or these waste materials are actually what we consider as valuable in many cases. Therefore, when we apply systems biology and engineering in this scope, we try to redirect these metabolic fluxes to maximize production. So this is what people try to do, genetic engineers, metabolic engineers, classically. But we need to be aware that in the metabolism, there is a competition for resources. So at the moment that we redirect these fluxes towards product, this means we take away resources, carbon flux, et cetera, towards biomass production.

And the challenge here is that this biomass is actually the catalyst of our system. So in this case, we may have higher yield, so we produce more product per substrate, but this reduces the overall volumetric productivity in bioreactors. So this is one example of a classical metabolic trade-off yield versus productivity in bioprocesses. Naturally, when people go in this direction, there is a lack of flexibility of adaptability to tune this during my process. This is, in a way, a static control approach. Furthermore, if I modify my cells too much or very intensively, I may lead into energy and redox imbalances or even resource burden, right? Because, for instance, if I overexpress a new metabolic pathway with many different new enzymes that need to be synthesized by the cell, this costs the cell energy, but also precursors, right?

So even though we have these different scopes or schools towards optimizing bioprocesses, they are not well integrated. And that's what I actually want to change with the research that we do in the group. But the second challenge that I see in biotechnology is that despite there being new bioproduction paradigms that aim to address these gaps that I have presented in the previous slide, there are still a lot of gaps for its implementations. So, for example, one concept that we focus a lot in the group is dynamic metabolic control through, in this case, metabolic cybergenetics. So the idea is, imagine that we have external inputs, for example, light intensity, that can regulate potentially the expression at the transcription level of proteins. This protein can be an enzyme, and by regulating the enzyme production, I can regulate naturally the metabolic flux that that enzyme catalyzes in metabolism. And in that way, I can have very specific nodes of regulation in the metabolic network at unprecedented levels.

I can also use light to manipulate, for example, whether a protein is active or not through appropriate photoreceptors. And as I said, this can give me the opportunity to tune metabolism in different ways. Ideally, we should be able also to get readouts of my metabolic system, so understand the state of my system such that I can take optimal decisions. These decisions in this context could be the light inputs, for example, that could readapt my system towards maximizing my process performance. So conceptually, this is very nice, right? However, there are questions that remain. One is, how can we steer metabolism optimally? So what are those optimal decisions that would keep my process operating robustly and efficiently?

Also, we need to be aware that this information that is captured by the cell is actually encoded by a lot of intracellular states, such that if I want to measure these intracellular states in real time, well, I would, in principle, need a lot of sensors that are actually not even available for the intracellular metabolism. So how can we estimate those metabolic states in real time then? Are there ways? There are, and that's what I will be also presenting later. And another idea that we need to take into account is, well, cell factories are very complex. So even though we can see a cell as kind of a black box, in reality, once we apply any input or new condition into the cell, that triggers a change of phenomena that actually will eventually give me a phenotype, a given process behavior. So if we actually zoom in there, we have different phenomena. So one is, for example, transcription from genes to mRNA.

Then we have translation from mRNA to proteins. Besides, we can also have post-translational modifications of those proteins to make them actually functional. And then these proteins would catalyze biochemical reactions in cells, and the amount of proteins, plus other things such as salt regulation, would also determine the upper bound of my metabolic fluxes. Right? So this is very complex. Besides, there are also feedback in cell regulation across all of these different aspects. Overall, if we were to define this mathematically, we are dealing with non-linear, stochastic, multi-rate, and multi-scale dynamics. From the stochastic point of view, one thing to notice is that this will lead to biological variability, and that is a common thing we see when we operate bioprocesses. So even though we may have the exact same setup, exact same strain, we may still have different process behavior because of this biological variability defined by a stochastic phenomenon that happens inside the cell.

Multi-rate, that's also an important thing from an engineering point of view, because once I apply an input, it may take some time before I see a change in the process behavior. So mathematically, you could define this as delays that also may complicate mathematical modeling and optimization. And besides that, cells are not operating as a single cell system. We are actually growing cells in bioreactors where we may have thousands of cells, as a big population of cells that are actually carrying out the process. Overall, this makes bioprocesses difficult to model, estimate, and control. Another piece of context of my research is bioprocess optimization remains still very heuristic.

So it is still very recipe-driven, trial-and-error driven, where operators are given sort of recipes, step one, step two, step three, and fingers crossed, the process will behave as expected. So it is kind of an open-loop framework, right? There are no corrective actions properly being taken in real time. In many cases, people would apply design of experiments, where they have some factors that they would basically change over different levels and try to find that optimal point. However, this framework is still kind of static, right? So you have some inputs and outputs, and you don't really focus on the transient behavior of my process. And bioprocesses are dynamic systems, so this kind of neglects this complexity, and therefore limits the potential optimization space that we are working with.

And also, if we have too many factors, for example, four or five, we are already into a situation where we would need to run too many experiments. For example, with five, with full factorial, we may need to do 32 experiments. That's too much, and of course, when I say too much, that's in terms of time and money. What people try to do in biotechnology, going a little bit more forward is they would try to split bioprocesses into stages. So one stage might be, well, what if I grow my cells first? For example, this can be controlled through an optogenetic signal. And later on, at some point, I switch to production.

So in this case, this would try to balance this trade-off I presented earlier between growth and production, right? And maybe, and this applies also to recipes. So once they find a correct set of inputs or maybe some switching points for my system, in this case, for a two-stage bioprocess, they would kind of lock this, and they would call it my kind of golden batch, and that would define my recipe. As I mentioned before, this typically operates in open loop, so no corrective actions are taken. Therefore, this may lead to moderate to poor reproducibility of my process, and it cannot adapt to uncertainties. And I also like to present this kind of figure or animation where I always say if, let's say, this is my entire optimization space and this is my true optima, and I'm aiming to minimize some function with the current approaches that we have, design of experiments, recipe-driven operation.

We, in the best case, may be, let's say, close to that real minima, but never really at that minima. It's really hard to get there given the limited scope of our optimization. So the strategy that we have in the group is what we call biotechnology systems engineering, which essentially is a systems of systems framework that integrates three main concepts. So systems biology, we want to understand the mechanistic behavior of the cell and metabolic networks. Process systems engineering, we are aware that the cell is not an isolated being, but they are usually operating in a bioreactor environment, so we need to understand the process as well.

And scientific machine learning tools, where we want to use the available physics of my biology and my processes towards building surrogate models, digital twins that can also capture data and can also help us reduce reality gaps between models and… and reality. So this is the approach we have. Now, what I will present next is some selected research topics such that you can get an idea of the specific research and capabilities that we have and we are developing in the group. So one mechanistic modeling approach we work is constraint-based modeling. So in constraint-based modeling, we use metabolic networks to understand cellular behavior. For example, in this simplified or small metabolic network, we can model the reactions that occur through different metabolic pathways in metabolism, and we can, for example, with these reactions, create mass balances.

We can build mass balances governing this metabolic network. So in these mass balances, one thing to notice is that we will have a stoichiometric matrix, but we will also have this flux vector. The flux vector are literally all of the flux or reaction rates across the metabolism. The problem with this is if we would just try to solve this metabolic network, we may end up having underdetermined systems. So we may have many more fluxes in this vector V than actually equations. So we end up with an underdetermined system. In constraint-based modeling, this is kind of circumvented by considering the cell as a self-optimizing automaton. So we assume the cell has an intrinsic objective.

It could be maximizing its growth rate due to evolutionary reasons, and we add biologically feasible or sound constraints. So besides the mass balances, we say, well, there may be transport capacity aspects. There are also enzyme capacity constraints. So if we don't have enzymes, there cannot be flux, and also we may consider that producing those enzymes towards enabling new reaction rates or the reaction rates in metabolism, they are not cheap. We need to invest ATP and precursors, and also density constraints, as the cell has a given density that it can afford. And once we have these mathematical models, we can basically answer the following question: What is that resulting metabolic flux distribution in the cell by solving this as an optimization problem?

Furthermore, if we know that we can tune metabolism through tunable inputs, for example, using optogenetics, using light to manipulate gene expression, we may also add this as further constraints in these optimization problems. That would, of course, constrain specific metabolic fluxes where we have actuation on. And we have studied this in the context of dynamic metabolic control. For example, how changing these inputs dynamically would affect my process performance, and how can I optimize, therefore, these dynamic trajectories of my input in real time. Examples, applications we have considered are, for instance, in E. coli, where we have worked on maximizing lactate production efficiency as well as ethanol production efficiency. In the first case, considering ATP, the energy state as our degree of freedom to tune metabolism through concepts such as ATP wasting, and also, in the second case of the ethanol fermentation from glycerol, in this case, we can consider inputs such as the oxygen uptake rate to enable transitions, let's say, very detailed transitions between aerobic and microaerobic states.

The second concept we are exploring in our research is surrogate metabolic modeling. So we are aware that these mechanistic models are actually quite complex, can be quite complex to solve, specifically when it comes to dynamic modeling approaches, and especially when we embed these models into model-based optimization and control strategies. So that is why we have been working on surrogate metabolic models, where, for example, we can learn the metabolic flux distribution under different tunable inputs and conditions via deep neural networks or neural networks or actually any machine learning regressor that is appropriate, such that we don't need to solve an optimization problem of the cell for each given condition, but we can actually just get immediately the response of the metabolic network given that input or constraint.

So, as I said, this is motivated by the fact that we want to link intracellular and extracellular metabolic domains for more tractable optimization and control. If we don't do this, if we actually embed the optimization problem that I presented before into a model-based optimization control strategy, we actually would be dealing with bi-level dynamic optimization schemes that also involves aspects such as game theory and so on that make things a bit not very easy to solve. So that's the motivation. Examples of these kind of approaches that we have worked are in itaconic acid production by E. coli, where we can manipulate flux through the tricarboxylic acid cycle, or also the ethanol production via manipulation of the acetate synthesis flux. Also to manipulate the ATP maximum capacity the cell can produce and thereby manipulating also metabolic pathways towards maximizing ethanol and balancing it with growth.

Another aspect that we explore is machine learning-supported hybrid modeling, where the main idea is how to fuse bioprocess knowledge with data-driven concepts. One example is recombinant protein production by CHO cells. So in this work, we had a macroscopic model that could model the dynamics of macroscopic states such as glucose, amino acids, lactate, biomass, and of course, my recombinant protein. However, there were still mismatches. So the available model structure couldn't really capture the full dynamics. Therefore, we, in this case, and that's the concept of hybrid modeling, we can augment this structure with machine learning components such as neural networks. That's what I did in order to have better prediction capabilities.

And another concept that is tangential to this research is we also want to, as I said, understand how intracellular and macroscopic domains are integrated in the bioreactor. So we use this information to further constrain dynamic metabolic flux analysis methodology, such that based on what was happening macroscopically, we could also propagate that through the metabolism and understand how key metabolic fluxes would be changing over time, leveraging, of course, the metabolic structure of my network, in this case of CHO cells. And another concept or aspect we are quite interested in is uncertainty estimation. So when we do modeling, we want to know how sure we are about the model that we are developing, not just giving any model, but how certain I am that's actually the right model. And this is a concept we also explored in this work.

Another example of hybrid modeling we have explored is in the context of recombinant protein production by Pichia pastoris with optogenetic induction. Without saying too many details, in this system, we have two degrees of freedom. So one is the EL222 copy number. This is my optogenetic transcription factor, and another degree of freedom is the light intensity at which I operate my system. So learning the dynamics of this complex system is not trivial because we have too many phenomena that may not be very well understood, given also the data that we have at hand. So in this case, what we applied is a Gaussian process, which is a nonparametric machine learning method to map parameters of my macroscopic model with basically these two different degrees of freedom that we have to operate our system. The gene copy number, which is an intracellular degree of freedom, and the light at which we shine our bioreactor dynamically, which is, let's say, our macroscopic degree of freedom.

And in that case, we were able actually to learn this complex behavior using data-driven approaches and model our system quite well. Another example I have here to present is, again, going in the direction of trying to link intracellular and extracellular domains. We had the question of how to link proteomic profiles, so proteomic data to predicted process dynamics. In this specific example, we had a library of single gene knockouts in Saccharomyces cerevisiae, and the issue here is, well, proteomic profiles is high dimensional. So we have high dimensional features and how to filter out those important features towards then using this lower dimension proteomic feature space towards linking that to macroscopic models. Well, that's what we were working here with, and we used Gaussian processes to…

Well, first of all, we used a concept such as permutation feature importance to isolate those important features, and then based on some dynamic experiments that we carried out, we were able to link those proteomic features in the reduced space with my dynamics through Gaussian processes. And another example of research that we are currently doing is multiscale mechanistic modeling. In this case, we are moving away, let's say, from hybrid approaches, and we really want, in this example, to understand the mechanistic behavior of my bioprocess. And this is relevant, especially when mechanistic understanding is relevant depending on the context of the project. So in this case, we are dealing with toxin-antitoxin systems inside synthetic microbial consortia towards polychromatic fermentations.

The motivation here is the following. If we have only one cell and we do metabolic genetic engineering in that cell only, we may burden the cell. So it's either we are adding new metabolic enzymes that the cell will need to synthesize, creating a resource burden. But also it might be that we need to delete a lot of genes, and this may disturb, for instance, the energy and redox balances of my cell, creating a cell that is not very fit for the process. So when we have this kind of problem, we can apply the concept of division of labor. So instead of giving all of the work to one population of cells, we can split that work among different microbes, and that's actually a very smart idea.

I always give the example, for example, when I give this to students, I tell them: imagine I give you homework, and I ask you to solve 100 exercises for tomorrow. Probably, you will get quite stressed out, and probably you won't solve them very well or efficiently. But what if instead, I split this homework into 100 students? So now each person will just need to do one exercise. Now, probably, very likely, the exercise will be better solved overall. So this is the same concept with consortia. The problem with consortia is that if there is no way to control the population of these cells, well, the fastest-growing population will, at some point, outgrow the other ones, and what was supposed to be a consortia now is a monoculture.

In collaboration with the Avalos Lab at Princeton University, we have been working on new strategies to use light intensity using these toxin-antitoxin systems to manipulate the growth rate of microorganisms. And we have been working on mathematical models that explain this specific optogenetic regulation and how to link that with the extracellular bioreactor dynamics. So this is the kind of example that I refer to when I say pure mechanistic modeling. Another example of mechanistic modeling, it is a project with CINVESTAV in Mexico. So we want to understand astaxanthin production dynamics with Pichia rhodosinum. This is a yeast under induced oxidative stress. So Pichia rhodosinum, when it is under oxidative stress, it causes that metabolic flux redistribution in the cell, given some, let's say, disturbance in the redox state of my system. So this can be represented, for example, through the ratio of the NADH/NAD⁺.

But this behavior causes that I will start producing, for example, astaxanthin due to the accumulation of reactive oxygen species. And all of these dynamics, so that explains the interaction between the intracellular redox state and the bioreactor dynamics and how to optimize this induced oxidative stress. That's part of the questions we try to answer with these mechanistic models that we are developing. Now, as I mentioned before, all of these modeling strategies that we follow, we do that because we want to make them useful for optimization and control. So, here I present in a more structured way what I mean by optimization and control. So imagine that we have a cost function.

This could be maximizing volumetric productivity, or I may have a minimization problem where I have some reference tracking or some trajectory that I want to track over time. It doesn't really matter. Let's consider it as an abstract cost function, and I can maximize or minimize this subject to the dynamic model. So this is basically our digital twin that we can embed there, plus additional constraints such as technical safety, economic constraints that may be relevant to the process. And if I solve this problem, I can obtain these inputs. So these are the inputs that I can apply in my process to fulfill this maximization problem, so to speak.

Now, if we do that without any feedback, so we just solve optimization, we apply the inputs. That's what we call an open-loop optimization. In many cases, that is what we can actually afford in experiments because of the lack of sensors for real-time monitoring. But as I said, if we actually could monitor the state in real time, then basically what we could do is to update the state of my system, the initial condition of my problem, so to speak, to recompute these input trajectories over time. We can do that, for instance, using model predictive control strategies. Another concept that we are exploring is batch-to-batch model adaptation. So let's say I cannot apply real-time feedback control, but I can gather data after each of my batches. So we have considered strategies such as reality gap closure through learning residuals using, for instance, Gaussian processes, where I can use these residuals, which is a data-driven model, to iteratively improve the knowledge of my system batch to batch.

And therefore, after, as we have demonstrated in publications, we could rapidly get to the real system dynamics after a few batches. And of course, this means that we can embed this also in adaptive model-based optimization strategies. When it comes to soft sensors, so as I said, if I actually want to be able to measure the state of my system, I need either hardware sensors, but in many cases, there are no available sensors, so I need to do something. And that is where soft sensors come into place. They are also called virtual sensors, where I aim to reconstruct the full state of my system from partial information. So from, let's say, easy to measure variables to reconstruct the system by inferring also the hard to measure variables or states, and with further capabilities such as noise filtering and model adaptation. An example that we have worked a lot is a moving horizon estimation.

So in moving horizon estimation, instead of solving an optimization towards the future to optimize my process, we solve an optimization towards a past horizon of measurements. To keep it simple, there, what we try to minimize is, let's say, the data, the past measurements with the predicted measurements or states from my model. And we can also consider things such as the arrival cost, or state noise that we can, as I said, include in our optimization problem, but that's a bit out of the scope of this presentation. And we have applied these strategies mainly towards estimating intracellular states such as proteins and metabolites in different works. So yeah, it has shown to be very efficient in these tasks.

And another research direction that we have is now machine learning-driven optimization, specifically using reinforcement learning for decision-making policies in bioprocesses. So in reinforcement learning, well, there are different flavors of reinforcement learning, but the one we are currently working with is mainly policy gradient methods. So in policy gradients, we still want to maximize some cost function. But now instead of considering my process as deterministic, we consider it as a naturally stochastic process, so a Markov decision process. Therefore, we are, in this case, now maximizing the expectation of that cost function. And the decision variable, so to speak, of this reinforcement learning problem is a policy pi.

The policy pi is a policy distribution where I could sample my inputs given the current state of my system. This could be parameterized, for example, using deep neural networks, where the state of my system would define a probabilistic distribution, for example, a Gaussian distribution, from which I could sample those inputs. And naturally, those parameters defining this deep neural network are updated, or we learn those parameters such that we can maximize this expectation of my cost function. In reinforcement learning, there are nice things to note. For example, what we call as controllers in classical control engineering, in this context, it will become my machine learning agent making those decisions. What we call as control inputs in a more control engineering context, here in reinforcement learning, we would call them actions. And the environment that the agent explores is actually our process in this context. And another thing I didn't mention, but it's also very relevant, the way in which we could inform the agent that we are making right or good decisions are via rewards that are basically mathematically encoded functions that would tell the agent how well we are doing by making given actions. And this is also very appropriate, and we have been using this approach when it comes to uncertain systems.

And as I said, biological systems are quite uncertain, quite stochastic. So it naturally matches that context of bioproduction. Furthermore, in many cases, the models I have are not tractable for model-based optimization and control. Therefore, we need to circumvent that, for example, using a model-free version such as reinforcement learning. And we have been working with, as I said, examples of reinforcement learning may include, in this context, the dynamic metabolic control, but also multi-set point and multi-trajectory tracking, for instance, when it comes to microbial consortia applications, and working with different types of inputs, like amplitude. The amplitude is literally a continuous input that you just vary from at different ranges, from a minimum to a maximum level, or even a pulse width modulation, which is very relevant to many bioprocesses, where the input is not applied just continuously, but actually it is a binary on/off action, on and off, that just changes when you move from the on to the off state. So it's more like a switching time problem instead.

For example, parameterized by duty cycles. And finally, everything that I have been showing you so far focuses on external control. So we have external inputs that, yes, can influence the intracellular metabolism, but the actuation always comes from the extracellular environment. However, we have been also working on in-cell controllers. So in in-cell controllers, we can use biochemical reaction networks that if you analyze them mathematically, they can carry properties of a PID-type controller. So we may have a proportional integral derivative and in more advanced setups, as we have explored in a recent publication, also acceleration gains of this type of PID, PIDA controllers. Now, these controllers intracellularly are, as I said, encoded via biochemical reaction networks.

We have proven also that we can achieve a good in-cell reference tracking given a good design of these biochemical reaction networks, even under stochastic conditions. So we have been working with these kind of concepts. As well as very recently, also biomolecular signal differentiators. Again, this is in collaboration with Princeton University. And in this context, what we were trying to answer is, well, what if I have a network, an intracellular network of interest, and its output is, in this case, U, but I cannot measure U directly? So in those contexts, what we can create are differentiators. So the differentiator is another kind of sensing unit in the cell, a network on its own, that instead of giving me the concentration of this, it will tell me how fast it is changing.

So it will tell me the speed or acceleration of this U state. And of course, this could be very useful for process monitoring, and these are also concepts we are currently exploring. And all of these, we are trying to focus not just on deterministic conditions, but actually more realistic biological conditions which are stochastic, as we have discussed previously. So with this, I hope I have given you a good overview of the motivation, the context, and the topics that, the direction that we are following as a group, the Systems Engineering for Biotechnology group at UCD. And I would like to finish with this slide. So I'm always happy to connect with the biotechnology bioprocess community to discuss challenges and create collaborations, and also see how we can translate many of these scientific findings into industrial applications.

So thank you very much. And finally, here I have some acknowledgement of the different institutions I have been working on throughout my career, but also many of the works that I presented had also these affiliations or co-affiliations. So, well, with this, thank you very much for your attention.

More from Invert University

Closed-System Cell Washing and Concentration for MSCs13:27

Closed-System Cell Washing and Concentration for MSCs

Beatriz Menéndez Yeves · Takeda

Jul 2026

The RNAbox: Continuous RNA Manufacturing42:01

The RNAbox: Continuous RNA Manufacturing

Prof. Zoltán Kis · University of Sheffield

Jul 2026

Read the paper
Automated Water-Free Thawing Device18:50

Automated Water-Free Thawing Device

Raúl Valero Carabias · Takeda / Universidad de Extremadura

Jul 2026

Read the paper