Best Object Detection Models to Run Locally in 2026: Table Transformer Structure Recognition

Rankings 2026-09-28 Last updated 2026-09-28 12 min read By Q4KM

Quick Answer

Microsoft’s Table Transformer Structure Recognition is the top pick as of September 2026 under the rule “newest release first, then downloads,” winning the release-date tie on downloads; recent releases are too few, so the newest available candidates are listed, with no public benchmark yet for any candidate [1][3][4][6]. In order, the ranking is Table Transformer Structure Recognition (Microsoft), Table Transformer Detection (Microsoft), DEtection TRansformer (DETR) with ResNet-50 (Facebook), and DEtection TRansformer (DETR) with ResNet-101 (Facebook) [1][3].

Key Takeaways

How do local object detection models compare on specifications and benchmark availability?

Model Org Params Quant/VRAM Released (date) License Key benchmark (date)
Table Transformer Structure Recognition [1] Microsoft [1] 29M [1] 4-bit weights: 14.5 MB estimated (params x 0.5 bytes); runtime VRAM: not published [1] 2022-10-14 [1] MIT [1] no public benchmark yet
Table Transformer Detection [3] Microsoft [3] 29M [3] 4-bit weights: 14.5 MB estimated (params x 0.5 bytes); runtime VRAM: not published [3] 2022-10-14 [3] MIT [3] no public benchmark yet
DEtection TRansformer (DETR) with ResNet-50 [4] Facebook [4] 42M [4] 4-bit weights: 21 MB estimated (params x 0.5 bytes); runtime VRAM: not published [4] 2022-03-02 [4] Apache-2.0 [4] no public benchmark yet
DEtection TRansformer (DETR) with ResNet-101 [6] Facebook [6] 61M [6] 4-bit weights: 30.5 MB estimated (params x 0.5 bytes); runtime VRAM: not published [6] 2022-03-02 [6] Apache-2.0 [6] no public benchmark yet

Which object detection models should you run locally?

1. Table Transformer Structure Recognition

Table Transformer Structure Recognition by Microsoft ranks first because its Hugging Face publication date ties the newest listed release, while its download count wins the tiebreak.[1][3][4][6] Ranking note: no public benchmark scores any candidate, so the order is newest release first, then downloads.[1][3][4][6] The category has too few recent releases; the newest available candidates are included.[1][3][4][6] Scope: this ranking admits only object detection models from labs with a published paper or leaderboard record, using Hugging Face’s object-detection pipeline tag.[1][2][3][4][5][6]

The model has 29M parameters, a listed context length of 1K tokens, and F32 weights occupying 0.1 GB.[1] Its 4-bit weight memory is approximately 14.5 MB, estimated (params x 0.5 bytes) from the cited 29M parameters.[1] Treat that calculation as a weight-storage budget when planning local hardware. The estimate does not establish total runtime memory, a minimum system-RAM requirement, or compatibility with a particular GPU. The license is MIT.[1]

Use the model for table structure recognition in a document-extraction workflow, consistent with Microsoft’s “PubTables-1M: Towards comprehensive table extraction from unstructured documents” paper.[1][2] The caveat is age and unverified comparative performance: its first Hugging Face publication was 2022-10-14, and its status here is “no public benchmark yet” as of 2026-09-25.[1] Its position reflects publication order and adoption; the ranking provides no measured accuracy or local inference-speed advantage over the other candidates.[1][3][4][6]

2. Table Transformer Detection

Table Transformer Detection by Microsoft ranks second under the ordering rule “newest release first, then downloads.”[1][3][4][6] Its Hugging Face publication date matches Microsoft’s Table Transformer Structure Recognition, but its download count is lower.[1][3] The ranking admits only object detection models from labs with a published paper or leaderboard record, using the Hugging Face object-detection pipeline tag. The eligible family has too few recent releases, so the newest available candidates are included; this model dates to October 14, 2022.[3]

The model has 29 million parameters, a listed context length of 1K tokens, F32 weights of 0.1 GB, and an MIT license.[3] For local hardware planning, 4-bit weight storage is approximately 14.5 MB, estimated (params x 0.5 bytes) from the cited 29 million parameters.[3] That calculation covers weights alone: runtime memory, image processing, and intermediate tensors require additional capacity. A specific GPU or system-RAM requirement cannot be established from the listed weight size alone.

Use Table Transformer Detection for the table-detection stage of a document extraction workflow, consistent with Microsoft’s “PubTables-1M: Towards comprehensive table extraction from unstructured documents” work.[2][3] The key caveat is “no public benchmark yet” within this ranking’s comparison. Its placement reflects publication date and download ordering, so the position does not establish superior detection accuracy or local inference speed.[1][3][4][6]

3. DEtection TRansformer (DETR) with ResNet-50

DEtection TRansformer (DETR) with ResNet-50 by Facebook ranks third among the listed candidates.[1][3][4][6] The ranking admits only object detection models from labs with a published paper or leaderboard record and the Hugging Face object-detection pipeline tag.[2][5] Too few releases qualify for the recency gate, so the newest available candidates are included.[1][3][4][6] Ordering is newest release first, then downloads; Microsoft’s Table Transformer Structure Recognition leads because it shares the newer publication date and has more downloads.[1][3][4][6]

Facebook lists 42M parameters, 0.2 GB of F32 weights, a context length of 1K tokens and an Apache-2.0 license.[4] Weight memory at 4-bit is 21 MB, estimated (params x 0.5 bytes) from the listed 42M parameters.[4] That calculation covers weights alone; runtime memory also needs room for image processing, intermediate tensors and the inference framework. A specific GPU or system RAM requirement cannot be established from the weight size alone.

General image object detection is the intended use, as described in Facebook’s “End-to-End Object Detection with Transformers.”[5] Hugging Face publication was 2022-03-02, with 349,004 downloads over the last 30 days as of 2026-09-25.[4] Downloads break the publication-date tie with Facebook’s DEtection TRansformer (DETR) with ResNet-101.[4][6] The performance caveat is no public benchmark yet in this ranking: its position does not establish an accuracy or speed advantage.

4. DEtection TRansformer (DETR) with ResNet-101

DEtection TRansformer (DETR) with ResNet-101 by Facebook ranks fourth in this selection.[6] The ordering is “newest release first, then downloads,” because no public benchmark scores the candidates against one another. Facebook published this checkpoint on Hugging Face on March 2, 2022; its 17,800 downloads over the preceding 30 days put it behind the tied-release ResNet-50 checkpoint as of September 25, 2026.[4][6] The selection has too few recent releases, so it includes older available models.[1][3][4][6]

The ranking admits only object detection models from labs with a published paper or leaderboard record, using Hugging Face’s object-detection pipeline tag. Facebook’s checkpoint has 61 million parameters, listed F32 weights of 0.2 GB, a listed context length of 1K tokens, and an Apache-2.0 license.[6] Its associated paper, “End-to-End Object Detection with Transformers,” was published on May 26, 2020.[5] Benchmark status: no public benchmark yet for this comparison; the paper date is not a benchmark date.

Consider this checkpoint for local object-detection experiments where the Apache-2.0 license suits your project.[6] For hardware budgeting, its 61 million parameters imply 30.5 MB of weight storage at 4-bit precision—estimated (params x 0.5 bytes).[6] That calculation covers weights alone; it does not establish working quantization support or total runtime memory. A specific GPU or system RAM requirement cannot be established from the listed specifications. The caveat is straightforward: placement here does not demonstrate an accuracy or speed advantage.

How can you estimate weight memory from parameter counts?

Estimate weight memory by multiplying parameter count by bytes per parameter: for 4-bit weights, use estimated (params x 0.5 bytes).[1][3][4][6] Treat the result as a weight-storage estimate, not a measured memory requirement for running detection.

Microsoft’s microsoft/table-transformer-structure-recognition has 29M parameters: 14.5 MB estimated (params x 0.5 bytes).[1] Microsoft’s microsoft/table-transformer-detection also has 29M parameters, giving 14.5 MB estimated (params x 0.5 bytes).[3] Equal parameter counts produce equal estimates under the same storage assumption.

Facebook’s facebook/detr-resnet-50 has 42M parameters: 21 MB estimated (params x 0.5 bytes).[4] Facebook’s facebook/detr-resnet-101 has 61M parameters: 30.5 MB estimated (params x 0.5 bytes).[6] Read those figures as arithmetic estimates for hypothetical quantized weights, rather than reported checkpoint sizes.

The listed checkpoints contain F32 weights, so the quantized estimates do not describe the supplied files.[1][3][4][6] A smaller calculated weight budget also does not establish that a compatible quantized checkpoint is available.

Use the calculation to compare weight-storage budgets before choosing an implementation. Keep any hardware-fit claim separate: the parameter calculation alone does not establish total runtime memory, available quantization support, or detection performance. A practical comparison should label estimated weight storage clearly and reserve measured memory figures for a documented runtime measurement.

Which licenses cover these object detection models?

Microsoft’s microsoft/table-transformer-structure-recognition and microsoft/table-transformer-detection use the MIT license;[1][3] Facebook’s facebook/detr-resnet-50 and facebook/detr-resnet-101 use the Apache-2.0 license.[4][6] Record the license alongside the exact checkpoint identifier when documenting a local deployment, rather than keeping a generic license entry for object detection.

Within Microsoft’s Table Transformer pair, the license stays MIT whether you select the structure-recognition checkpoint or the detection checkpoint.[1][3] Both models accompany Microsoft’s paper “PubTables-1M: Towards comprehensive table extraction from unstructured documents.”[2] Keep their model identifiers separate in your dependency inventory even though their license labels match.

Within Facebook’s Detection Transformer (DETR) pair, the license stays Apache-2.0 across the listed checkpoints.[4][6] Both accompany the paper “End-to-End Object Detection with Transformers.”[5] Record whichever checkpoint you actually distribute or deploy, together with its license documentation.

For deployment review, use those checkpoint-specific license labels as the starting point. Review the applicable license text before modifying or redistributing weights, and check the terms for accompanying code and dependencies separately. A model’s license entry should remain attached to that model in your records; avoid treating it as a blanket license for the surrounding application.

How should you choose a model when no public benchmark compares the candidates?

Choose by release recency, then downloads as a tiebreak, and use license and weight requirements to narrow your deployment choice; the ordering does not establish detection accuracy.

The ranking admits only object detection models from labs with a published paper or leaderboard record, using Hugging Face’s object-detection pipeline tag.[2][5] The ranking note is: newest release first, then downloads. The candidate pool has too few releases within the requested twelve-month window, so the newest available candidates are listed despite their older publication dates.[1][3][4][6]

Microsoft’s microsoft/table-transformer-structure-recognition ranks first because it shares the latest Hugging Face publication date, October 14, 2022, with Microsoft’s microsoft/table-transformer-detection and has more downloads.[1][3] Facebook’s facebook/detr-resnet-50 follows, then Facebook’s facebook/detr-resnet-101; both were first published on March 2, 2022, and downloads break their tie.[4][6] The established picks are the structure-recognition model, table-detection model and DETR ResNet-50, the three most-downloaded candidates; their recency exemption does not itself confer first place.[1][3][4]

For every candidate, the benchmark status is “no public benchmark yet.” Downloads indicate adoption, not comparative accuracy. Evaluate candidates on representative local images before committing.

License and weight storage provide concrete deployment filters. Both Microsoft models use MIT licensing and have 29M parameters each: 4-bit weight storage is estimated (params x 0.5 bytes) at 14.5 MB each.[1][3] Both Facebook models use Apache-2.0 licensing.[4][6] A weight-storage estimate alone does not establish runtime memory requirements or hardware fit.

Frequently Asked Questions

Which object detection model ranks first for local use?

Microsoft’s Table Transformer for Structure Recognition (microsoft/table-transformer-structure-recognition) ranks first because its publication date ties with Microsoft’s Table Transformer for Detection (microsoft/table-transformer-detection), while its download count is higher.[1][3] The detection variant follows, then Facebook’s Detection Transformer with ResNet-50 (facebook/detr-resnet-50), then Facebook’s Detection Transformer with ResNet-101 (facebook/detr-resnet-101).[3][4][6] The ordering reflects publication dates and downloads; no public benchmark yet establishes an accuracy winner among these candidates.

How is the ranking determined?

Ranking note: this ranking admits only object detection models from labs with a published paper or leaderboard record (Hugging Face pipeline tags: object-detection).[1][2][3][4][5][6] No public benchmark scores any candidate, so the ordering is newest release first, then downloads. The newest available candidates fall outside the release window.[1][3][4][6] Established picks are microsoft/table-transformer-structure-recognition, microsoft/table-transformer-detection and facebook/detr-resnet-50, the most-downloaded entries; their recency exemptions cannot confer first place.[1][3][4][6]

How much memory would quantized weights require?

For either Table Transformer variant, 29M parameters imply 14.5 MB of 4-bit weights, estimated (params x 0.5 bytes).[1][3] DETR with ResNet-50 has 42M parameters: 21 MB estimated (params x 0.5 bytes).[4] DETR with ResNet-101 has 61M parameters: 30.5 MB estimated (params x 0.5 bytes).[6] Those calculations cover weights alone and do not establish runtime memory requirements or quantization compatibility.

What context length and original weight sizes are listed?

Each candidate lists a context length of 1K tokens.[1][3][4][6] Both Table Transformer variants list F32 weights of 0.1 GB, while both DETR variants list F32 weights of 0.2 GB.[1][3][4][6] Treat the context value as listed metadata: the supplied specification does not establish an image-resolution limit or a document-page limit. The weight sizes likewise do not establish total inference memory.

What licenses apply to these local object detection models?

Both Microsoft Table Transformer variants carry the MIT license.[1][3] Both Facebook DETR variants carry the Apache-2.0 license.[4][6] The license therefore stays the same when choosing between the listed variants within either family. For an engineering inventory, record the exact repository identifier alongside its license so the selected checkpoint remains identifiable when deployment requirements or model choices change.

Are these recent releases with public benchmark results?

Neither family’s listed checkpoints are recent: both Table Transformer entries were first published on Hugging Face on October 14, 2022, and both DETR entries on March 2, 2022.[1][3][4][6] The benchmark status for every candidate is “no public benchmark yet”; no benchmark score or evaluation date is available for this comparison. Paper publication dates establish publication history, not comparative detection performance.[2][5]

Sources

  1. microsoft/table-transformer-structure-recognition model card (Hugging Face) — 2026-09-25
  2. PubTables-1M: Towards comprehensive table extraction from unstructured documents — 2021-09-30
  3. microsoft/table-transformer-detection model card (Hugging Face) — 2026-09-25
  4. facebook/detr-resnet-50 model card (Hugging Face) — 2026-09-25
  5. End-to-End Object Detection with Transformers — 2020-05-26
  6. facebook/detr-resnet-101 model card (Hugging Face) — 2026-09-25

Get these models on a hard drive

Skip the downloads. Browse our catalog of 985+ commercially-licensed AI models, available pre-loaded on high-speed drives.

Browse Model Catalog