.. _using_abc_vs_protocols: Using abc.ABC vs Protocols ========================== There are many opportunities within programming in general to solve problems by creating a standard/framework/mechanism that generalises those problems so they can be solved consistently. This collection of static typing advice originated in a codebase that used ``pkgutil.resolve_name`` in many places as a form of dependency injection for "manager" objects that abstract away implementation details depending on how the software was configured at runtime. This created a need to think about how we can use this mechanism in a type safe way. The advice here remains specific to that use case, but can be useful more generally. Prelude: Safety at Runtime vs Statictime ---------------------------------------- An important concept to consider when designing an interface is what and when the interface is providing safety. Safety in this context refers to the idea that when we write code our linters and tests will accurately and completely tell us when we are doing something that isn't guaranteed to work at runtime. And when at runtime, our code will stop and rollback before data is corrupted. At runtime we have ABC classes that provide a level of runtime protection, which is explained below. Within the context of static type checking this concept of safety boils down to "does this variable name a value having this attribute and does the value I get from that attribute in turn have other attributes I depend on and so on". The type checker runs at static time, which means we are talking about what is and isn't *guaranteed*, as opposed to what is or isn't *true at runtime*. To illustrate this point, take this example: .. code-block:: python class A: def __init__(self) -> None: self.some_information = 1 def modifies_a(a: A, *, attr: str) -> None: setattr(a, attr, 2) my_variable = A() modifies_a(my_variable, attr="sneaky_other_information") print(my_variable.sneaky_other_information) In this example, we can easily read it and see this will work at runtime. But the static type checker doesn't read code the way a human typically does, and at any kind of scale it's essentially an NP complete problem for it to understand whether this kind of code would work at runtime (it's also quite smelly code!) There are definitely limitations to the current state of Python type checking, but it does offer a lot of protection and often times, when it complains, it's mypy's love language for saying the flow of information needs to be adjusted. Prelude: How do ``abc.ABC`` classes work? ----------------------------------------- Python offers "Abstract Base Classes" under the `abc standard library `_ that offer a way to formalise what methods/attributes an object should have, such that we have runtime protection from making objects that don't have these methods/attributes. The way it works is you subclass `abc.ABC `_ and mark required attributes with the ``@abc.abstractmethod`` decorator. For example: .. code-block:: python import abc class BaseProcessor(abc.ABC): @property @abc.abstractmethod def some_information(self) -> SomeInformation: ... @abc.abstractmethod def process(self) -> None: ... A Protocol version of this would look like .. code-block:: python from typing import Protocol class BaseProcessor(Protocol): @property def some_information(self) -> SomeInformation: ... def process(self) -> None: ... There are a few key differences between the two however: - An ABC class can also contain an implementation, whereas a protocol can not - An ABC class represents a runtime concept whereas the protocol represents a static time only concept. - An ABC class can inherit from other concrete objects, whereas it often doesn't make sense for a protocol to do that - You can use ``isinstance`` checks with an ABC class, whereas doing that with a Protocol doesn't make sense and the ``@typing.runtime_checkable`` decorator doesn't do what it looks like and is best avoided. The nice thing about the abc class is that we can only instantiate instances of the abstract class that have an implementation for these attributes marked with ``abc.abstractmethod``. For example .. code-block:: python import attrs @attrs.frozen class ConcreteProcessor(BaseProcessor): some_information: SomeInformation def process(self) -> None: print(self.some_information) When we instantiate this ``ConcreteProcessor``, the metaclass that this class is created with (given we inherit from ``abc.ABC``) will make sure that none of the attributes on the instance are marked as abstract. If our subclass overrides all the abstract attributes on the base class, no exception will be raised. Whereas: .. code-block:: python import attrs @attrs.frozen class ConcreteProcessor(BaseProcessor): some_information: SomeInformation When this implementation is instantiated, the ``.process`` will be the abstract method from the parent class and an exception will be raised complaining that we have not implemented all abstract instances. This is how we get runtime safety, because we are guaranteed that if we have an instance of this class, then all required attributes have some implementation. And we combined this with mypy enforcing Liskov substitution principle to ensure that the types of these attributes are compatible with what we expect. Prelude: Breaking safety ------------------------ It's very easy to break the safety these constructs give us! In both cases it's via code that should be better anyways! To break the safety for an abc class is to do something like this: .. code-block:: python import attrs @attrs.frozen class ConcreteProcessor(BaseProcessor): some_information: SomeInformation def process(self) -> None: raise NotImplementedError In this case we can instantiate the object because it satisfies the checks that ABC makes, but we don't actually have an implementation we can rely on! In this case we don't actually have an implementation and this can happen when a base class has too many responsibilities and the codebase ends up with with subclasses that only need to implement some of the methods. The better approach is to break up the functionality provided into multiple base classes rather than having one object with too many responsibilities. To break a Protocol means ignoring the type checker, for example: .. code-block:: python import pkgutil class Manager(protocol): def do_something(self) -> None: ... def get_manager() -> Manager: return pkgutil.resolve_name(settings.MY_MANAGER) The return of ``resolve_name`` is an ``Any`` and nothing in this example guarantees that it matches the protocol that is in the return annotation of the function. When we use the function however we don't know that the function is lying! .. code-block:: python get_manager().do_something() # mypy would think this is fine, but at runtime nothing guarantees that! And the goal at the end of the day with all our linters is to find and fix bugs before users that depend on our deployed code finds them! Patterns: It is useful to provide a common ABC with no restrictions ------------------------------------------------------------------- Specifically when it has been decided to make a standard interface using an ABC class, start with a class that makes no assumption about how it's implemented! This could be called the "Common" interface. Even if there is only one default implementation, it makes it far easier to work with the interface when requirements change enough in the future that a second default implementation becomes useful. It's much easier to add behaviour to a class than it is to remove behaviour from a class! .. code-block:: python # The interface that is being provided class CommonManager(abc.ABC): @abc.abstractmethod def functionality_one(self) -> str: ... @abc.abstractmethod def functionality_two(self) -> str: ... @abc.abstractmethod def functionality_three(self) -> str: ... # A default implementation in terms of `SomeObj` that is created using attrs @attrs.frozen class BaseManager[T_Obj: SomeObj](CommonManager, abc.ABC): @property @abc.abstractmethod def obj(self) -> T_Obj: ... @abc.abstractmethod def functionality_one(self) -> str: ... @abc.abstractmethod def functionality_two(self) -> str: ... def functionality_three(self) -> str: if self._some_choice_for_functionality_three(): self.obj.do_something() @abc.abstractmethod def _some_choice_for_functionality_three(self) -> bool: ... # A concrete implementation using our default BaseManager for `ElecObj` @attrs.frozen class ElecManager(BaseManager[ElecObj]): obj: ElecObj def functionality_one(self) -> str: return obj.one def functionality_two(self) -> str: return "stuff" def _some_choice_for_functionality_three(self) -> bool: return True # A concrete implementation using our default BaseManager for `GasObj` @attrs.frozen class GasManager(BaseManager[GasObj]): obj: GasObj def functionality_one(self) -> str: return obj.make_one() def functionality_two(self) -> str: return "tree" def _some_choice_for_functionality_three(self) -> bool: return False This example [attempts to] demonstrate the idea that what the object provides is not always a 1:1 match with how it provides it. In this case the implementation for ``ElecManager`` still provides ``.functionality_three`` but implements it in terms of ``_some_choice_for_functionality_three`` By having the ``CommonManager`` ancestor only be concerned with the functionality a consumer should know about we don't limit other implementations of this object with concerns related to this family of ``SomeObj`` objects that may not be relevant in all circumstances. The lesson here is separation of concern: Specifically separating what is being provided from how it is provided. Patterns: Prefer to not use runtime_checkable --------------------------------------------- The python ``typing`` module provides a decorator by the name ``runtime_checkable`` that makes it so that you can make a Protocol class work with the ``isinstance`` builtin. The functionality provided by this decorator however is not how it seems: .. code-block:: python from typing import runtime_checkable, Protocol # Don't do this @runtime_checkable class MyProtocol(Protocol): def do_something(self, one: int, *, two: str) -> bool: ... class MyImplementation: def do_something(self) -> dict[str, int]: return {} def takes_something(instance: object) -> None: assert isinstance(instance, MyProtocol) # mypy believes that if that doesn't raise an exception then instance # must match MyProtocol and the following will work instance.do_something(1, two="two") # The assertion type narrowing in the function won't fail # Because runtime_checkable doesn't check types # It only checks existance. takes_something(MyImplementation()) The problem with the ``runtime_checkable`` is that it doesn't check that the types of those attributes are correct. And especially for generic protocols, it's not really possible for it to do that. The alternative is to check against known classes that implement the desired protocol. This is because type narrowing is fundamentally a runtime check because it's at runtime that we know what the object actually is. At static time each choice is only a possibility. And so we must check the object against runtime concepts rather than static time concepts to decide what the object actually is. Patterns: type narrowing dynamically retrieved logic ---------------------------------------------------- A place where Protocols begin to lose their type safety is when we retrieve logic that is typed as ``typing.Any``. For example the standard library `pkgutils.resolve_name `_ which takes a string and returns the result of doing a python import with that value: .. code-block:: python from django.conf import settings import pkgutil from typing import Protocol class Bike(Protocol): def steer(self, direction: str) -> ... def get_bicycle() -> Bike: return pkgutil.resolve_name(settings.BIKE_CLASS)() In this case ``resolve_name`` can return anything, and some unsuspecting code may try to do a ``get_bicycle()`` and it may take a few layers before the code fails because that object wasn't actually a bike! It is much safer to instead do something like: .. code-block:: python from django.conf import settings import pkgutil import abc class CommonBike(abc.ABC): @abc.abstractmethod def steer(self, direction: str) -> ... def get_bicycle() -> CommonBike: constructor = pkgutil.resolve_name(settings.BIKE_CLASS) instance = constructor() assert isinstance(instance, CommonBike) return instance .. note:: ❗ Prefer ``assert isinstance`` over ``assert issubclass`` so any callable object can be found instead of only class objects. In this pattern we assume that whatever we get from ``resolve_name`` is a callable object and returns an instance of the object we want. Patterns: Using ``resolve_name`` in a type safe way for objects that do take extra information ---------------------------------------------------------------------------------------------- Let's increase the difficulty a little from the previous section! Sometimes it's useful to provide information when we get our object: .. code-block:: python def get_bicycle(color: str) -> CommonBike: constructor = pkgutil.resolve_name(settings.BIKE_CLASS) # But constructor is Any, how do we know it takes color!?!?! instance = constructor(color=color) assert isinstance(instance, CommonBike) return instance In this case we want an object that we know takes in ``color`` as a keyword argument but we have an ``Any`` and we don't want to limit ourselves to what kind of factory we get. .. note:: 😅 fun fact, a class constructor is a factory that returns an instance of that class. One way to make this type safe is to expect an intermediate object that is used to create what we want. This involves a little more boilerplate, but it's worth it for code that ages so much better as requirements change over time. .. code-block:: python import abc import pkgutil class BikeMaker(abc.ABC): def make(self, *, color: str) -> CommonBike: ... def get_bicycle(color: str) -> CommonBike: constructor = pkgutil.resolve_name(settings.BIKE_CLASS_MAKER) maker = constructor() assert isinstance(maker, BikeMaker) return maker.make(color=color) Note that with this pattern it becomes ok for the ``make`` method to have a Protocol as a return annotation because once we do our type narrowing against an abc class we know we have an object that is returning something matching that protocol (assuming that implementation isn't doing something smelly like a type ignore comment!). The key takeaway is a great pattern that universally applies is to assume the return of ``resolve_name`` is a callable object that returns what we want. The variation here is that what we want from ``resolve_name`` is an object that can then in turn make the final object we are returning. Patterns: Using ``resolve_name`` in a type safe way for generic objects ----------------------------------------------------------------------- We can take the previous section and add one more thing onto it to enable us to return generic objects: .. code-block:: python import abc import pkgutil class CommonLaptop[T_CPU: CommonCPU](abc.ABC): @property @abstractmethod def cpu(self) -> T_CPU: ... @abc.abstractmethod def run_code(self, code: str) -> None: ... class LaptopMaker(abc.ABC): def make[T_CPU: CommonCPU]( self, *, cpu_type: type[T_CPU] ) -> CommonLaptop[T_CPU]: ... def get_laptop(cpu_type: type[T_CPU]) -> CommonLaptop[T_CPU]: constructor = pkgutil.resolve_name(settings.LAPTOP_MAKER) maker = constructor() assert isinstance(maker, LaptopMaker) return maker.make(cpu_type=cpu_type) def get_arm_laptop() -> CommonLaptop[ARMCPU]: return get_laptop(ARMCPU) def get_x86_laptop() -> CommonLaptop[X86]: return get_laptop(X86) The trick to understanding this is that a ``TypeVar`` is the same as any other type annotation except that it's the caller that decides what the type annotation is rather than the definition. And so to get a generic object with a specific type, we need to provide information about what that inner type is, which in turn means we need a factory that takes in information and is able to generate our final object. .. note:: 😁 Note a really nice feature from this pattern of assuming ``resolve_name`` returns a callable object is that it gives us a really easy built in way of avoiding import time side effects, because being a callable means we can provide code that is only run when we execute the object inside our get function! Patterns: Prefer managers that *have* objects rather than *take* objects ------------------------------------------------------------------------ Let's say we have this manager class: .. code-block:: python class BaseManager: def get_vendor_name_from_dispenser(self, dispenser: CommonDispenser) -> str: return dispenser.vendor_name This is very restrictive because it means when we try to specialise such a manager managers we need to worry about all the different dispenser types rather than only the specialised one. There are two solutions to this 1. Make CommonManager generic to the type of the dispenser 2. Make the CommonManager not care about the dispenser Let's talk about option (2) .. code-block:: python import pkgutil class CommonDispenser """ A placeholder for this example that represents a generalized dispensing mechanism that specialized dispensers can subclass. """ class CommonManager(abc.ABC): @abc.abstractmethod def do_something_with_dispenser(self) -> str: ... @attrs.frozen class CandyManager(CommonManager): dispenser: SpecialisedCandyDispenser def do_something_with_dispenser(self) -> str: self.dispenser.some_candy_feature() return "treats" @attrs.frozen class SnackManager(CommonManager): dispenser: SpecialisedSnackDispenser def do_something_with_dispenser(self) -> str: self.dispenser.some_snack_feature() return "bag of snacks" class MakeManagerFromDispenser(abc.ABC): def make(self, *, dispenser: CommonDispenser) -> CommonManager: ... def get_manager_from_dispenser(self, *, dispenser: CommonDispenser) -> CommonManager: constructor = pkgutil.resolve_name(settings.MANAGER_FROM_DISPENSER_MAKER) maker = constructor() assert isinstance(maker, MakeManagerFromDispenser) return maker.make(dispenser=dispenser) class SpecialisedMakeManagerFromDispenser(abc.ABC): def make(self, *, dispenser: CommonDispenser) -> CommonManager: match dispenser: case SpecialisedCandyDispenser(): return CandyManager(dispenser=dispenser) case SpecialisedSnackDispenser(): return SnackManager(dispenser=dispenser) case _: raise UnknownDispenserType The difference is that we make a manager instance per instance of the dispense. Rather than one manager instance that may take multiple dispensers. It also means that the interface we work against is no longer generic and doesn't leak where the information is coming from (because the interface does not expose the dispenser itself). It also means we can do things like this: .. code-block:: python class CandyManager(CommonManager): dispenser: SpecialisedCandyDispenser def get_candy_manager(self, dispenser: SpecialisedCandyDispenser) -> CandyManager: return CandyManager(dispenser=dispenser) And use the manager we've already created to create other managers that get the objects we are working with. This makes it much easier to cater interesting edge cases where we need to pass around and build up context to get the right functionality for the situation. Patterns: AdHoc Protocols are best defined where they are used -------------------------------------------------------------- Whilst there can be a case for defining protocols in a common place, it's likely to follow the same logic in Go and define protocols (which are like Go interfaces) where they are used. And use a protocol as an expectation of what is required rather than a contract of what is provided. For example: .. code-block:: python from typing import Protocol class HasProcessMethod(Protocol): def process(self) -> bool: ... def do_processing(self, processor: HasProcessMethod) -> bool: return processor.process() Doing this means any object that has a ``process`` method that doesn't require any arguments and returns a boolean can be used by this function. This function is the perfect lego piece, it makes no assumption about what is passed in or its purpose other than what it wants to do with the object. It depends on the caller to know what purpose is being satisfied by using its functionality. It can be a smell for this kind of function to have opinions about what higher purpose it has! It is but a soldier doing its one assigned task! The main exception to leaning towards abc classes for protocols is for representing callable objects. In most cases this replaces using ``collections.abc.Callable`` For example, instead of: .. code-block:: python from collections.abc import Callable def make_processor() -> Callable[[], None]: return lambda: None Instead have: .. code-block:: python from typing import Protocol class Processor(Protocol): def __call__(self) -> None: ... def make_processor() -> Processor: return lambda: None There are two reasons for this: - Means the Callable has a name - Allows us to represent complex signatures that have positional and keyword only arguments and star args. It also makes a lot of sense to represent a constructor as a separate protocol. For example: .. code-block:: python class MyObj(Protocol): @property def one(self) -> str: ... @property def two(self) -> int: ... class MyObjConstructor(Protocol): def __call__(self, *, one: str, two: int) -> MyObj: ... This is because there are many ways to create an object and the ``__init__`` method is only one of those ways. And often times it is useful to split what we construct from where we get the data to construct it from. For example, this statically works: .. code-block:: python @attrs.frozen class MyImplementation: one: str two: int three: bool def takes_constructor( constructor: MyObjConstructor, *, one: str, two: int ) -> MyObj: return constructor(one=one, two=two) takes_constructor(functools.partial(MyImplementation, three=True), one="one", two=1) But this does not .. code-block:: python @attrs.frozen class MyImplementation: one: str two: int three: bool def takes_constructor( constructor: type[MyObj], *, one: str, two: int ) -> MyObj: return constructor(one=one, two=two) takes_constructor(functools.partial(MyImplementation, three=True), one="one", two=1)