EngGPT 2-16B-A3B e' il Large Language Model italiano sviluppato da Engineering Group per ottimizzare i processi di business garantendo controllo, efficienza e trasparenza nell'adozione dell'AI, in conformita' all'AI ACT.
EngGPT 2-16B-A3B e' il Large Language Model italiano sviluppato da Engineering Group per ottimizzare i processi di business garantendo controllo, efficienza e trasparenza nell'adozione dell'AI. Sviluppato from scratch in Italia, senza utilizzare modelli preesistenti, garantisce trasparenza sui processi di training e conformita' nativa all'AI Act. Progettato per contesti enterprise e PA, grazie all'architettura Mixture-of-Experts ottimizza costi di training e di inferenza garantendo prestazioni elevate, risultati accurati e ampia possibilita' di personalizzazione sul proprio dominio specifico. E' rilasciato open weight per analisi indipendenti e test, con piena disclosure del processo di training all'interno del technical report. Include la capacita' di doppio ragionamento e risposta del modello: una specifica ed una rapida, in lingue diverse in base all'esigenza. Nativamente ottimizzato per la lingua italiana, si integra nei sistemi esistenti per accelerare casi d'uso come assistenza knowledge-intensive, automazione di workflow complessi, analisi di dati complessi in NL, interpretazione di schemi industriali. E' agent-ready per orchestrare flussi agentic e automazioni su processi e use case mission-critical. Disponibile come servizio su marketplace AWS per un accesso immediato, con standard di sicurezza certificati e scalabilita' globale.
Open Weight & Auditability - il modello e' pubblicato su HuggingFace, con pesi, documentazione e benchmark pubblici. Anche il processo di training e' trasparente: dataset, pipeline, tecniche, metriche e relative limitazioni sono disponibili nel "Technical Report" suArxiv.
Training and Inference efficiency - Architettura Mixture-of-Experts e double reasoning consentono output di qualita' elevata e costi sotto controllo. Agent-ready per orchestrare e ottimizzare i processi di business.
Customizability - Costi di training ridotti grazie al MoE consentono una rapida ed efficace verticalizzazione sul proprio dominio. Ottimizzato per la lingua e il contesto italiano, e' disponibile nelle 6 principali lingue europee (Italiano, Inglese, Tedesco, Spagnolo, Portoghese e Francese).
AWS Marketplace now accepts line of credit payments through the PNC Vendor Finance program. This program is available to select AWS customers in the US, excluding NV, NC, ND, TN, & VT.
You pay by the hour for each host running model inference, so cost scales with instance size and how long you run it. Pricing splits across three instance families: ml.g4dn.12xlarge, several ml.g5 sizes, and a range of ml.g6e sizes from xlarge up to 48xlarge. The ml.g4dn.12xlarge and ml.g5 instances each offer two modes: batch, for grouped jobs, and real-time, for live requests. The ml.g6e instances are real-time only. Larger instances carry higher hourly rates. You choose the family, size, and mode that fit your workload.
Top-of-mind questions for buyers
What is the difference between the batch and real-time inference modes for billing?
Both meter running host-hours, but they serve different workload types. Batch mode groups many requests into scheduled jobs, suited to bulk processing. Real-time mode keeps a host live to answer requests as they arrive. You choose the mode that matches your workload, and you pay for each running host by the hour.
Am I charged when an inference host is stopped or idle?
Charges accrue per running host-hour. A host that is fully stopped does not accrue software charges. If you keep a real-time host live to answer requests, it bills for the whole time it runs, whether or not requests arrive. Underlying AWS infrastructure fees may still apply separately.
What does one host-hour cover for running this model?
One host-hour is one hour of one running instance hosting the model for inference. The model uses a Mixture-of-Experts design with about 16 billion total parameters and 3 billion active per token. You pick an instance family and size, and each running host bills by the hour.
www.eng.it
Helpful?
Vendor refund policy
No refund allowed
How can we make this page better?
Tell us how we can improve this page, or report an issue with this product.
Give us feedbackReport a problem with this product or seller
Legal
Vendor terms and conditions
Upon subscribing to this product, you must acknowledge and agree to the terms and conditions outlined in the vendor's End User License Agreement (EULA).
Content disclaimer
Vendors are responsible for their product descriptions and other product content. AWS does not warrant that vendors' product descriptions or other product content are accurate, complete, reliable, current, or error-free.
An Amazon SageMaker model package is a pre-trained machine learning model ready to use without additional training. Use the model package to create a model on Amazon SageMaker for real-time inference or batch processing. Amazon SageMaker is a fully managed platform for building, training, and deploying machine learning models at scale.
Deploy the model on Amazon SageMaker AI using the following options:
Real-time inference
Deploy the model as an API endpoint for your applications. When you send data to the endpoint, SageMaker processes it and returns results by API response. The endpoint runs continuously until you delete it. You're billed for software and SageMaker infrastructure costs while the endpoint runs. AWS Marketplace models don't support Amazon SageMaker Asynchronous Inference. For more information, see Deploy models for real-time inference .
Batch transform
Deploy the model to process batches of data stored in Amazon Simple Storage Service (Amazon S3). SageMaker runs the job, processes your data, and returns results to Amazon S3. When complete, SageMaker stops the model. You're billed for software and SageMaker infrastructure costs only during the batch job. Duration depends on your model, instance type, and dataset size. AWS Marketplace models don't support Amazon SageMaker Asynchronous Inference. For more information, see Batch transform for inference with Amazon SageMaker AI .
Version release notes
Primo rilascio del modello LLM deployato su AWS SageMaker AI
Supporto inferenza real-time e batch, scalabile
Pronto per integrazione pipeline e applicazioni cloud-native
AWS Support is a one-on-one, fast-response support channel that is staffed 24x7x365 with experienced and technical support engineers. The service helps customers of all sizes and technical abilities to successfully utilize the products and features provided by Amazon Web Services.
Be the first to review this product. We've partnered with PeerSpot to gather customer feedback. You can share your experience by writing or recording a review, or scheduling a call with a PeerSpot analyst.