Building a Neural Processor (NPU) similar to TPU, to achieve hardware acceleration for running local LLMs