AWS NewsDropbox EngineeringGitHub BlogGoogle DevelopersMeta EngineeringNetflix TechBlogStripe Engineering

Total Articles: 17 from 7 sources


AWS News

1. Runtime instances: persistent compute for production AI agents on Amazon Bedrock AgentCore

URL: https://aws.amazon.com/blogs/aws/runtime-instances-persistent-compute-for-production-ai-agents-on-amazon-bedrock-agentcore/

Published: 2026-08-06 22:58

Summary:

Announcing runtime instances in Amazon Bedrock AgentCore—persistent, managed EC2 infrastructure for production AI agents with multi-agent collaboration, GPU support, and sessions lasting up to 14 days.


2. Amazon DynamoDB now supports real-time vector search at any scale

URL: https://aws.amazon.com/blogs/aws/amazon-dynamodb-now-supports-real-time-vector-search-at-any-scale/

Published: 2026-08-05 14:45

Summary:

DynamoDB now supports native vector search with single-digit millisecond latency at 99%+ recall It is designed for any scale, even trillions of vectors and requires zero infrastructure management.


3. AWS Weekly Roundup: Price reduction of GPT models in Bedrock, CloudWatch managed collectors for Prometheus metrics, and more (August 3, 2026)

URL: https://aws.amazon.com/blogs/aws/aws-weekly-roundup-price-reduction-of-gpt-models-in-bedrock-cloudwatch-managed-collectors-for-prometheus-metrics-and-more-august-3-2026/

Published: 2026-08-03 16:12

Summary:

Last week I had the joy of participating in Amazon’s “Bring Your Kids to Work Day” with my 7 year old son We commuted together into the New York City office, his first real rush hour train ride, and spent the day exploring how Amazon uses AI, machine learning, and robotics to deliver packages to […]

Dropbox Engineering

1. How our universal content processing platform Riviera evolved for AI and beyond

URL: https://dropbox.tech/infrastructure/how-our-universal-content-processing-platform-riviera-evolved-for-ai-and-beyond

Published: 2026-07-20 15:00

Summary:

Riviera is the Dropbox content processing platform that’s been iteratively improving content transformation in our products for roughly a decade.

GitHub Blog

1. A guide to slash commands in the GitHub Copilot app

URL: https://github.blog/ai-and-ml/github-copilot/a-guide-to-slash-commands-in-the-github-copilot-app/

Published: 2026-08-06 19:49

Summary:

Go beyond chat in the GitHub Copilot app with these slash commands They’ll help you plan, collaborate, automate, and customize your dev workflow The post A guide to slash commands in the GitHub Copilot app appeared first on The GitHub Blog.


2. How we took malware advisories beyond npm

URL: https://github.blog/security/supply-chain-security/how-we-took-malware-advisories-beyond-npm/

Published: 2026-08-06 16:51

Summary:

GitHub malware advisories no longer stop at npm Here’s how we wired OpenSSF’s malicious-packages data into the Advisory Database, and why we built the pipeline paranoid The post How we took malware advisories beyond npm appeared first on The GitHub Blog.


URL: https://github.blog/ai-and-ml/github-copilot/how-the-github-legal-team-used-copilot-cli-to-streamline-their-workflows/

Published: 2026-08-04 19:02

Summary:

Learn how to build tools to simplify how you work—without writing a single line of code The post How the GitHub legal team used Copilot CLI to streamline their workflows appeared first on The GitHub Blog.

Google Developers

1. ML Development in VS Code with Google Cloud Power: Workbench Extension Now Available

URL: https://developers.googleblog.com/ml-development-in-vs-code-with-google-cloud-power-workbench-extension-now-available/

Published: 2026-08-07 09:58

Summary:

The Google Cloud Workbench Notebooks extension for VS Code has officially launched, allowing developers to connect their local IDE to scalable, cloud-based Jupyter environments This integration streamlines the machine learning lifecycle by eliminating context switching and providing direct access to high-performance Google Cloud infrastructure To support transparency and community-driven innovation, the newly released extension is fully open-sourced and available on GitHub and the VS Code Marketplace.


2. Why we built ADK 2.0

URL: https://developers.googleblog.com/why-we-built-adk-20/

Published: 2026-08-07 09:58

Summary:

Answering the questions of “why we built ADK 2.0” This explains the rationale, some of the features, and why a developer should consider upgrading This will be published the day after ADK go 2.0 launches.


3. Build agentic full-stack apps with Genkit

URL: https://developers.googleblog.com/build-agentic-full-stack-apps-with-genkit/

Published: 2026-08-07 09:58

Summary:

The open-source Genkit framework has introduced the Agents API, a full-stack tool designed to simplify the complex plumbing of conversational AI by packaging message history, tool loops, and streaming into a single interface The API supports flexible, server- or client-managed state persistence—allowing for advanced workflows like history branching, long-running detached tasks, and multi-agent coordination—while seamlessly connecting backends to frontends via a unified wire protocol Currently available in preview for TypeScript and Go, it also integrates with the Genkit Developer UI to allow developers to easily test, debug, and inspect agent snapshots without writing client code.

Meta Engineering

1. From User Sequences to Scaling Laws: A Multi-Stage Architecture for Meta’s Ads Ranking

URL: https://engineering.fb.com/2026/08/05/ml-applications/from-user-sequences-to-scaling-laws-a-multi-stage-architecture-for-metas-ads-ranking/

Published: 2026-08-05 19:20

Summary:

Every day, Meta’s recommendation platforms handle billions of user interactions, generating rich temporal signals that capture individual preferences and intent across products, ads, and content In our 2024 post on sequence learning for ads recommendations, we showed how modeling the order and timing of user actions (rather than relying on static, manually engineered sparse features) […] Read More The post From User Sequences to Scaling Laws: A Multi-Stage Architecture for Meta’s Ads Ranking appeared first on Engineering at Meta.


2. GEM Training: How Meta Doubled the Efficiency of Its LLM-Scale Ads Foundation Model

URL: https://engineering.fb.com/2026/08/03/ml-applications/training-gem-at-llm-scale-meta-ads-recommendation-foundation-model/

Published: 2026-08-03 18:00

Summary:

Meta’s Generative Ads Recommendation Model (GEM), the foundation model behind ads recommendations across Instagram and Facebook, now trains at LLM scale on several thousand of the latest-generation GPUs This post goes into the details on how we achieved: doubling end-to-end (E2E) training efficiency to 20–25% Model FLOPs Utilization (MFU) while scaling training FLOPs 4x in […] Read More The post GEM Training: How Meta Doubled the Efficiency of Its LLM-Scale Ads Foundation Model appeared first on Engineering at Meta.


3. Exploring Hierarchical Interest Representation For Meta Ads Deep Funnel Optimization

URL: https://engineering.fb.com/2026/07/15/ai-research/exploring-hierarchical-interest-representation-for-meta-ads-deep-funnel-optimization/

Published: 2026-07-15 17:00

Summary:

Hierarchical Interest Representation is a research area for Meta Ads The innovations in Hierarchical Interest Representation are […] Read More The post Exploring Hierarchical Interest Representation For Meta Ads Deep Funnel Optimization appeared first on Engineering at Meta.

Netflix TechBlog

1. Modeling Device Capabilities for Analytics

URL: https://netflixtechblog.com/modeling-device-capabilities-for-analytics-e7607acebde8?source=rss----2615bd06b42e---4

Published: 2026-07-31 16:01

Summary:

To ensure the best possible user experience, we rely on a deep understanding of device capabilities We use a cumulative table to process information about the device’s capabilities By relying on data-driven insights, we can make informed decisions about which features to enable on specific devices, ensuring both performance and reliability.Modeling Device Capabilities for Analytics was originally published in Netflix TechBlog on Medium, where people are continuing the conversation by highlighting and responding to this story.


2. GenRec: Towards LLM-Native Recommendation at Netflix

URL: https://netflixtechblog.com/genrec-towards-llm-native-recommendation-at-netflix-f20be6f643e3?source=rss----2615bd06b42e---4

Published: 2026-07-30 20:10

Summary:

Raw logs of user history, item metadata, and context are transformed via context engineering into natural-language prompts and fed into the GenRec, which runs on vLLM in prefill-only mode and outputs scores for each catalog item, yielding a recommendation ranking.At a high level, GenRec:Verbalizes user histories, item metadata, and context as text.Post‑trains a Netflix‑adapted foundation LLM for ranking.Adds a catalog‑aware scoring head over Netflix titles.Uses reward signals to align with long‑term member value and business goals.Runs in prefill‑only mode on Netflix’s LLM serving stack for cost efficiency.In a large‑scale A/B test against a well‑tuned production ranker, GenRec achieves statistically significant improvements in both short‑term and long‑term online metrics, while using only a small fraction of the Phase‑2 labeled data and input signals As we increased Phase‑2 training data and enriched the input signals, GenRec’s offline metrics continued to improve.Online, we ran a large A/B test on batch‑compute recommendation surfaces, covering ~10% of Netflix traffic over ~4 weeks By verbalizing user histories, context, and item metadata, adding a catalog‑aware ranking head, using reward‑weighted objectives aligned to long‑term satisfaction and business goals, and serving efficiently on our LLM infrastructure, we obtain a model that improves on a strong production ranker while using far fewer Phase‑2 labels and input signals.GenRec is an early but promising step toward a more LLM‑centric recommendation stack at Netflix


3. In-House LLM Serving at Netflix

URL: https://netflixtechblog.com/in-house-llm-serving-at-netflix-a5a8e799ea2c?source=rss----2615bd06b42e---4

Published: 2026-07-17 21:32

Summary:

Serving Architecture OverviewDesign Decisions and ImplementationFour decisions shape this platform — engine, packaging, API surface, and rollout — presented in dependency order, since each one constrains the next.vLLM as the Paved-Path EngineThe platform was originally built on TensorRT-LLM, a performant inference engine at the time and already integrated with Triton — the compute backend in use within MSS.By summer 2025, two things had shifted: open-source engines had largely closed the performance gap with specialized stacks, and our workload mix had broadened to include embedding generation, prefill-only inference for ranking and retrieval, autoregressive decoding, and custom models with non-trivial per-step constraint logic The platform has to pin compatible versions when baking the service image, and prevent model authors from overriding the vLLM version at packaging time.Custom model logic We git-subtreed and patched the frontend to translate response_format into vLLM’s guided decoding parameters at request time.Deployment StrategiesWith API surface and engine in place, the question that remains is how new versions roll out without dropping requests

Stripe Engineering

1. Analyzing the evidence that helps businesses win “product not received” disputes

URL: https://stripe.com/blog/analyzing-the-evidence-that-helps-businesses-win-product-not-received-disputes

Published: 2026-07-21 00:00

Summary:

To understand what can influence win rates, we analyzed evidence packets from one million disputes over a 16-week period Here’s what the data shows and what it means for how you mitigate disputes.


Generated on 2026-08-07 09:58:20