Best Masked Language Models for Apple Silicon Macs in 2026: LFM2.5-Encoder-350M

Rankings 2026-09-27 Last updated 2026-09-27 14 min read By Q4KM

Quick Answer

LFM2.5-Encoder-350M (LiquidAI) is the top pick for Apple Silicon Macs as of September 2026 because it shares the newest release date and wins the downloads tiebreak; all candidates have no public benchmark yet, eligibility requires fill-mask models from labs with a published paper or leaderboard record, and the limited recent selection requires including the newest available older releases [1][2][3][4][5][6][7]. In order, the ranking is LFM2.5-Encoder-350M (LiquidAI), LFM2.5-Encoder-230M (LiquidAI), ESM-2 (esm2_t33_650M_UR50D) (facebook), ESM-2 (esm2_t6_8M_UR50D) (facebook), mDeBERTa-v3-base (microsoft), and DeBERTa-v3-base (microsoft) [1][2].

Key Takeaways

How do masked language models compare on memory, context, licenses and published benchmarks?

Model Org Params Quant/VRAM Released (date) License Key benchmark (date)
LFM2.5-Encoder-350M [6] LiquidAI [6] 354M [6] 4-bit weights: 177 MB estimated (params x 0.5 bytes) [6] 2026-07-27 [6] lfm1.0 [6] no public benchmark yet
LFM2.5-Encoder-230M [7] LiquidAI [7] 230M [7] 4-bit weights: 115 MB estimated (params x 0.5 bytes) [7] 2026-07-27 [7] lfm1.0 [7] no public benchmark yet
ESM-2 (esm2_t33_650M_UR50D) [4] facebook [4] 652M [4] 4-bit weights: 326 MB estimated (params x 0.5 bytes) [4] 2022-09-27 [4] MIT [4] no public benchmark yet
ESM-2 (esm2_t6_8M_UR50D) [5] facebook [5] 8M [5] 4-bit weights: 4 MB estimated (params x 0.5 bytes) [5] 2022-09-26 [5] MIT [5] no public benchmark yet
mDeBERTa-v3-base [1] microsoft [1] — — 2022-03-02 [1] MIT [1] no public benchmark yet
DeBERTa-v3-base [3] microsoft [3] — — 2022-03-02 [3] MIT [3] no public benchmark yet

Which masked language models should you consider for an Apple Silicon Mac?

1. LFM2.5-Encoder-350M

LFM2.5-Encoder-350M by LiquidAI ranks first because its July 27, 2026 release ties with LiquidAI’s LFM2.5-Encoder-230M, while its 29,446 monthly downloads exceed its sibling’s 7,111, breaking the recency tie.[6][7] The model has 354M parameters, a 128K-token context window and an lfm1.0 license.[6] Its benchmark status is “no public benchmark yet”; the ranking therefore does not establish superior prediction quality.

For local Apple Silicon work, consider it for fill-mask tasks where the advertised context window is useful.[6] Weight storage at 4-bit is approximately 177 MB, estimated (params x 0.5 bytes) from the cited 354M parameters.[6] The caveat is deployment: that calculation establishes neither an available quantized implementation nor total unified-memory requirements. Choosing a Mac requires validating runtime compatibility and memory use at the intended input length; a specific RAM recommendation cannot be established from weight storage alone.

Ranking note: this ranking admits only fill-mask models from labs with a published paper or leaderboard record. With no public benchmark scoring any candidate, the ordering is “newest release first, then downloads.” Too few recent releases qualify, so the newest available family releases are included. Established picks Microsoft mDeBERTa-v3-base,[1] Microsoft DeBERTa-v3-base,[3] and Facebook esm2_t33_650M_UR50D[4] receive recency exemptions based on downloads, follow the same ordering rule, and cannot take first place through that exemption.

2. LFM2.5-Encoder-230M

LFM2.5-Encoder-230M by LiquidAI ranks second because it shares the July 27, 2026 release date of LiquidAI’s LFM2.5-Encoder-350M but has fewer downloads: 7,111 versus 29,446 over the reported 30-day window.[6][7] The ordering is newest release first, then downloads; no public benchmark scores any candidate. Its benchmark status is “no public benchmark yet,” so its position should not be read as a demonstrated quality or Apple Silicon performance advantage.

The model has 230 million parameters, a context length of 128K tokens, and published F32 weights totaling 0.9 GB.[7] For local hardware planning, quantized weight storage would be 115 MB at 4-bit, estimated (params x 0.5 bytes) from the cited parameter count.[7] Actual unified-memory requirements also include runtime allocations and working memory. A specific Mac configuration cannot be recommended from weight storage alone, particularly when planning to use the advertised context length.

Consider the model for local fill-mask evaluation when long input support matters to your application. Validate your intended workload and context length in a compatible Apple Silicon runtime before committing to deployment. The practical caveat is licensing: LiquidAI distributes the model under lfm1.0.[7] Check those terms against your intended use before integrating the weights into a product.

3. ESM-2 (esm2_t33_650M_UR50D)

ESM-2 (esm2_t33_650M_UR50D) by facebook ranks third under the ordering rule: newest release first, then downloads.[4][6][7] An established pick, it was first published on Hugging Face on September 27, 2022, and recorded 2,063,734 downloads in the reporting window ending September 25, 2026.[4] Its position reflects release order, with downloads used only to break ties; no public benchmark yet establishes its performance against the other candidates. The older release remains included because the eligible selection has too few recent models.

The model has 652 million parameters, a context length of 1,026 tokens, and an MIT license.[4] Published F32 weights occupy 2.6 GB.[4] For hardware planning, its 4-bit weight storage is approximately 326 MB, estimated (params x 0.5 bytes) from the cited parameter count.[4] That calculation covers weights alone. A local Apple Silicon deployment also needs memory for execution, so the estimate does not establish a minimum Mac RAM configuration or confirm a compatible quantized runtime.

Choose this entry for local fill-mask evaluation when its MIT license and 1,026-token context meet your requirements.[4] The practical caveat is deployment uncertainty: no Apple Silicon throughput measurement or validated local memory requirement accompanies the listed specifications. Confirm runtime compatibility and measure memory use on your intended workload before selecting hardware.

4. ESM-2 (esm2_t6_8M_UR50D)

ESM-2 (esm2_t6_8M_UR50D) by facebook ranks fourth under the ordering rule: newest release first, then downloads.[4][5][6][7] Its status is “no public benchmark yet,” so the position does not establish comparative accuracy. Published on Hugging Face on September 26, 2022, the model follows the later releases ahead of it.[5] The shortlist includes older available models because too few recent releases qualify. Eligibility is limited to fill-mask models from labs with a published paper or leaderboard record.

The model has 8M parameters, a context length of 1,026 tokens, and an MIT license.[5] For an Apple Silicon Mac, its 4-bit weight storage would be approximately 4 MB, estimated (params x 0.5 bytes) from the cited 8M parameters.[5] Treat that figure as a weight-storage budget, rather than a total memory requirement. A local setup also needs a compatible runtime and memory for execution; a specific Mac configuration cannot be established from that estimate alone.

Consider it for local fill-mask experiments where a small parameter count matters to your storage budget.[5] The practical caveat is that estimated quantized weight storage does not demonstrate working quantization or execution on Apple Silicon. Verify runtime compatibility before choosing it for a local workflow, and evaluate task quality directly rather than treating its ranking as a performance result.

5. mDeBERTa-v3-base

mDeBERTa-v3-base by microsoft ranks fifth as an established pick, with an MIT license and a context length of 512 tokens.[1] The ordering is newest release first, then downloads; its publication date of March 2, 2022 places it behind the LiquidAI and facebook entries.[1][4][5][6][7] Within the same-date microsoft pair, its 5,455,850 downloads over the last 30 days put it ahead of microsoft’s DeBERTa-v3-base, with 2,797,571 downloads.[1][3]

The ranking admits only fill-mask models from labs with a published paper or leaderboard record. The family has too few recent releases, so the newest available entries are included alongside established picks. mDeBERTa-v3-base qualifies as an established pick through downloads, but popularity does not demonstrate model quality.[1][3][4][5][6][7] Its benchmark status is “no public benchmark yet.” The associated paper is “DeBERTa: Decoding-enhanced BERT with Disentangled Attention”; that paper’s presence does not establish comparative performance for this checkpoint.[2]

For an Apple Silicon Mac, consider this checkpoint for a local fill-mask evaluation where the MIT license and 512-token context meet your requirements.[1] Hardware sizing remains unverified: a defensible quantized-weight memory estimate requires a parameter count, and no specific chip or unified-memory capacity can be recommended here. Check runtime support and measure memory with your intended workload before committing. The practical caveat is the context ceiling: keep each input within the documented 512-token limit.[1]

6. DeBERTa-v3-base

DeBERTa-v3-base by microsoft ranks here as an established pick, with no public benchmark yet. Its first Hugging Face publication was March 2, 2022, the same date as microsoft’s mDeBERTa-v3-base.[3][1] The ordering is newest release first, then downloads. DeBERTa-v3-base follows its multilingual sibling because their release dates tie and its monthly downloads are lower: 2,797,571 versus 5,455,850 as of September 25, 2026.[3][1] Download counts resolve that tie; they do not establish model quality.

The ranking admits only fill-mask models from labs with a published paper or leaderboard record. With too few recent releases available, the list includes older models; DeBERTa-v3-base qualifies as an established pick through download popularity.[3] The model has a 512-token context limit and an MIT license.[3] Consider it for short-input fill-mask evaluation when that context budget meets your needs and MIT licensing suits your project. The context limit is the practical caveat: longer inputs require a separate handling strategy.

For local use on an Apple Silicon Mac, no defensible RAM requirement or quantized weight-memory estimate can be stated from the available specifications. A parameter count is needed to calculate a weight-only estimate, and that estimate would not establish total runtime memory. Before committing to deployment, verify that your chosen runtime supports the model and its intended task, then measure memory use with representative inputs. The ranking does not establish tested Mac compatibility or inference speed.

How do you estimate quantized weight memory for a masked language model?

Estimate quantized weight memory by multiplying the published parameter count by the storage per parameter: for a four-bit calculation, use estimated (params x 0.5 bytes).[6][7] Report the result as a weight-only estimate, and keep the parameter count beside it so readers can check the arithmetic.

LiquidAI’s LFM2.5-Encoder-350M has 354 million parameters: 177 MB estimated (params x 0.5 bytes).[6] LiquidAI’s LFM2.5-Encoder-230M has 230 million parameters: 115 MB estimated (params x 0.5 bytes).[7] Both calculations use decimal megabytes and assume every parameter receives the stated storage allocation.

Apply the same calculation to facebook’s esm2_t33_650M_UR50D: its published 652 million parameters give 326 MB estimated (params x 0.5 bytes).[4] Calculate from the parameter count when available; rounding in a published file size can make a file-size-based estimate less precise.

Treat each result as a starting point for an Apple Silicon memory budget. Quantization metadata, runtime buffers and memory used during inference sit outside the parameter-only calculation. Actual memory use also depends on the implementation and workload. Check the intended runtime and quantized artifact before making a fit claim; weight arithmetic alone does not establish local compatibility or total unified-memory requirements.

Which masked language models use the MIT license?

Microsoft’s mDeBERTa-v3-base (microsoft) [1] and DeBERTa-v3-base (microsoft) [3], plus Facebook’s facebook/esm2_t33_650M_UR50D [4] and facebook/esm2_t6_8M_UR50D [5], use the MIT license.

Microsoft’s models each have a context length of 512 tokens. [1][3] Both were first published on Hugging Face on March 2, 2022. [1][3] Neither entry provides a parameter count here, so a weight-memory estimate would require additional information. Their shared license and context length do not establish equivalent task performance.

Facebook’s esm2_t33_650M_UR50D has 652M parameters and a context length of 1,026 tokens. [4] Its 4-bit weight storage is approximately 326 MB, estimated (params x 0.5 bytes) from the cited 652M parameters. [4] Facebook’s esm2_t6_8M_UR50D has 8M parameters and the same 1,026-token context. [5] Its 4-bit weight storage is approximately 4 MB, estimated (params x 0.5 bytes) from the cited 8M parameters. [5] Those calculations estimate weight storage; they do not establish total runtime memory or verified Apple Silicon compatibility.

LiquidAI’s LFM2.5-Encoder-350M (LiquidAI) [6] and LFM2.5-Encoder-230M (LiquidAI) [7] instead use the lfm1.0 license. An MIT-only shortlist therefore excludes both LiquidAI entries. For a local deployment, treat license eligibility, task suitability and runtime support as separate checks before choosing a model.

How are newer releases and established picks ranked without public benchmark scores?

Newer releases and established picks are ranked by newest release first, then downloads, because no public benchmark scores any candidate.[1][3][4][5][6][7] The ranking admits only fill-mask models from labs with a published paper or leaderboard record, using Hugging Face’s fill-mask pipeline tag. The category has too few releases within the last twelve months, so the newest available candidates appear alongside older releases.[1][3][4][5][6][7]

LiquidAI’s LFM2.5-Encoder-350M takes the first position because it shares the newest publication date with LiquidAI’s LFM2.5-Encoder-230M and wins the downloads tiebreak.[6][7] Both were published on July 27, 2026; their respective downloads were 29,446 and 7,111 over the thirty days measured on September 25, 2026.[6][7] Both carry the status “no public benchmark yet.”[6][7] Popularity resolves their date tie; it does not demonstrate better model quality.

Facebook’s esm2_t33_650M_UR50D follows, then Facebook’s esm2_t6_8M_UR50D, published on September 27 and September 26, 2022, respectively.[4][5] Microsoft’s mdeberta-v3-base precedes Microsoft’s deberta-v3-base: both were published on March 2, 2022, and downloads break their tie.[1][3] Each of these older candidates also carries “no public benchmark yet.”[1][3][4][5]

The established picks are mdeberta-v3-base, deberta-v3-base and esm2_t33_650M_UR50D—the three download leaders in this candidate set.[1][3][4][5][6][7] Their exemption permits inclusion beyond the recency window, but cannot award the first position. Ordering follows publication date and the downloads tiebreak, never parameter size. For Apple Silicon buyers, the resulting order is a selection rule, not evidence of measured Mac performance.

Frequently Asked Questions

Which masked language model should I consider first?

LiquidAI’s LFM2.5-Encoder-350M ranks first because its July 27, 2026 release ties the newest candidate date, and its 29,446 downloads exceed its sibling’s 7,111 downloads over the reported period.[6][7] Recent releases are too few, so the ranking includes the newest available candidates alongside older releases.[1][3][4][5][6][7] The model has no public benchmark yet; its position does not establish superior performance on Apple Silicon.

How is the ranking ordered?

Ranking note: newest release first, then downloads.[1][3][4][5][6][7] Order: LFM2.5-Encoder-350M (LiquidAI) [6], LFM2.5-Encoder-230M (LiquidAI) [7], facebook/esm2_t33_650M_UR50D (Facebook) [4], facebook/esm2_t6_8M_UR50D (Facebook) [5], mDeBERTa-v3-base (microsoft) [1], DeBERTa-v3-base (microsoft) [3]. The ranking admits only fill-mask models from labs with a published paper or leaderboard record. Established picks—mdeberta-v3-base [1], deberta-v3-base [3] and esm2_t33_650M_UR50D [4]—bypass recency but cannot take first through that exemption. Every candidate has no public benchmark yet.

How much memory should I budget for quantized weights?

For 4-bit weights, LFM2.5-Encoder-350M’s 354M parameters imply 177 MB, estimated (params x 0.5 bytes).[6] LFM2.5-Encoder-230M’s 230M parameters imply 115 MB, estimated (params x 0.5 bytes).[7] The corresponding estimates are 326 MB for esm2_t33_650M_UR50D’s 652M parameters, estimated (params x 0.5 bytes),[4] and 4 MB for esm2_t6_8M_UR50D’s 8M parameters, estimated (params x 0.5 bytes).[5] Treat these as weight-only calculations, not measured Mac memory requirements.

Which context limits matter when choosing a model?

LFM2.5-Encoder-350M and LFM2.5-Encoder-230M each list a context length of 128K tokens.[6][7] The ESM candidates each list 1,026 tokens,[4][5] while mdeberta-v3-base and deberta-v3-base each list 512 tokens.[1][3] Match the listed limit to your intended input length. Treat a context specification as a selection criterion; do not interpret it as a measured latency or memory result for your Apple Silicon Mac.

Which candidates use the MIT license?

Microsoft’s mdeberta-v3-base and deberta-v3-base use the MIT license.[1][3] Facebook’s esm2_t33_650M_UR50D and esm2_t6_8M_UR50D also use MIT.[4][5] LiquidAI’s LFM2.5-Encoder-350M and LFM2.5-Encoder-230M instead list lfm1.0.[6][7] If MIT is a project requirement, select candidates from the Microsoft and Facebook entries. Review lfm1.0 separately before choosing a LiquidAI model for your intended use.

Does this ranking tell me which model runs faster on my Mac?

No public benchmark yet applies to every candidate here, so the ordering does not establish an Apple Silicon speed winner. The published context limits and licenses help narrow your choices,[1][3][4][5][6][7] but leave Mac runtime performance unresolved. Before adopting a model, verify that your chosen runtime supports its architecture and intended quantization, then measure memory use and latency with your actual inputs.

Sources

  1. microsoft/mdeberta-v3-base model card (Hugging Face) — 2026-09-25
  2. DeBERTa: Decoding-enhanced BERT with Disentangled Attention — 2020-06-05
  3. microsoft/deberta-v3-base model card (Hugging Face) — 2026-09-25
  4. facebook/esm2_t33_650M_UR50D model card (Hugging Face) — 2026-09-25
  5. facebook/esm2_t6_8M_UR50D model card (Hugging Face) — 2026-09-25
  6. LiquidAI/LFM2.5-Encoder-350M model card (Hugging Face) — 2026-09-25
  7. LiquidAI/LFM2.5-Encoder-230M model card (Hugging Face) — 2026-09-25

Get these models on a hard drive

Skip the downloads. Browse our catalog of 985+ commercially-licensed AI models, available pre-loaded on high-speed drives.

Browse Model Catalog