description Phi-3.5-MoE Overview
Phi-3.5-MoE is a mixture-of-experts large language model released by Microsoft in August 2024 as part of the Phi-3.5 family. The architecture uses 16 expert modules with 3.8 billion parameters each, of which 2 are activated per token, yielding 6.6 billion active parameters within a total of 42 billion. Microsoft reported that the model achieved competitive reasoning and coding benchmarks relative to larger dense models at substantially lower inference cost. The weights were released under an MIT license for research and commercial use.
help Phi-3.5-MoE FAQ
How many experts are in Microsoft's Phi-3.5-MoE model?
The model utilizes a mixture-of-experts architecture featuring a total of 16 experts. However, only a subset of these experts is activated for any given token during inference.
How many active parameters does Phi-3.5-MoE use during generation?
Despite having a large total parameter count, the model operates with roughly 6.6 billion active parameters during generation. This routing technique allows it to maintain strong reasoning while keeping compute costs low.
When did Microsoft release the Phi-3.5-MoE language model?
Microsoft released the Phi-3.5-MoE language model in 2024. It was part of their ongoing initiative to create highly efficient, small-footprint AI models.
What is the primary advantage of Microsoft's Phi-3.5-MoE architecture?
The primary advantage is its ability to offer strong reasoning capabilities relative to its low compute cost. By only activating a fraction of its total parameters at a time, it delivers high performance on standard hardware.
explore Explore More
Similar to Phi-3.5-MoE
See all arrow_forwardReviews & Comments
Write a Review
Be the first to review
Share your thoughts with the community and help others make better decisions.