You could use the plain api libraries for each llm and ipython notebooks, conceptually each block could be a node or link in the prompt chain, and input/output of each block is printable and visible to check which part of the chain is the part that is failing or has sub optimal outputs.