SDK_CLI_test - #3623
SDK_CLI_test#3623v-rkasula (raghu-microsoft) wants to merge 80 commits into
Conversation
* Add grpo job example * Replace client info with generics * Add missing copyright * newline at end of aml_setup.py * Fix black formatting issues and reduce dataset * Update sdk/python/jobs/grpo/src/grpo_trainer_rewards.py Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com> * Update sdk/python/jobs/grpo/src/BldDemo_Reasoning_Train.py Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com> * Update sdk/python/jobs/grpo/src/BldDemo_Reasoning_Train.py Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com> * Fix duplicated cell * Add to CODEOWNERS * Fix agenda image and dataset string * Replace demo with example * Update sdk/python/jobs/grpo/launch_grpo_command_job-med-mcqa-commented.ipynb Co-authored-by: Gayatri Penumetsa <181455625+gpenumetsa-msft@users.noreply.github.com> * Update sdk/python/jobs/grpo/launch_grpo_command_job-med-mcqa-commented.ipynb Co-authored-by: Gayatri Penumetsa <181455625+gpenumetsa-msft@users.noreply.github.com> * Update sdk/python/jobs/grpo/launch_grpo_command_job-med-mcqa-commented.ipynb Co-authored-by: Gayatri Penumetsa <181455625+gpenumetsa-msft@users.noreply.github.com> * Update sdk/python/jobs/grpo/launch_grpo_command_job-med-mcqa-commented.ipynb Co-authored-by: Gayatri Penumetsa <181455625+gpenumetsa-msft@users.noreply.github.com> * Address deployment related comments * Add README * Update sdk/python/jobs/grpo/launch_grpo_command_job-med-mcqa-commented.ipynb Co-authored-by: Gayatri Penumetsa <181455625+gpenumetsa-msft@users.noreply.github.com> * Remove duplicated model info section * Update sdk/python/jobs/grpo/launch_grpo_command_job-med-mcqa-commented.ipynb Co-authored-by: Gayatri Penumetsa <181455625+gpenumetsa-msft@users.noreply.github.com> * Update sdk/python/jobs/grpo/launch_grpo_command_job-med-mcqa-commented.ipynb Co-authored-by: Gayatri Penumetsa <181455625+gpenumetsa-msft@users.noreply.github.com> * Change dataset description --------- Co-authored-by: Sharvin Jondhale <shjondhale@microsoft.com> Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com> Co-authored-by: Gayatri Penumetsa <181455625+gpenumetsa-msft@users.noreply.github.com>
Co-authored-by: Mohd Javed Ansari <moansa@microsoft.com>
Add yeshwanth and harsha to code owners for GRPO
Changed Agenda to 'Plan of Action'
* Followup improvements to GRPO job example * Add more docs to the readme * Update sdk/python/jobs/grpo/README.md Co-authored-by: Gayatri Penumetsa <181455625+gpenumetsa-msft@users.noreply.github.com> * Update sdk/python/jobs/grpo/README.md Co-authored-by: Gayatri Penumetsa <181455625+gpenumetsa-msft@users.noreply.github.com> * Update README.md --------- Co-authored-by: Sharvin Jondhale <shjondhale@microsoft.com> Co-authored-by: Gayatri Penumetsa <181455625+gpenumetsa-msft@users.noreply.github.com>
Co-authored-by: Mohd Javed Ansari <moansa@microsoft.com>
Co-authored-by: Mohd Javed Ansari <moansa@microsoft.com>
* fail sdk installation workflow temporarily
* Replace azureml-defaults with azureml-inference-server-http==1.4.0 in environment configuration files * update azureml-inference-server-http * update inference-schema and joblib in conda.yaml * Replace azureml-defaults with azureml-inference-server-http==1.4.0 in environment configuration files * update azureml-inference-server-http * update inference-schema and joblib in conda.yaml --------- Co-authored-by: Salman Arshad <v-saarshad@microsoft.com>
* Adding mcp server content * updating content * add more content to it * adding more details * adding pr details * formating * updating readme file * updating readme * formating the content * formatting * applying formatng * applying formating * applying formatting * adding more ocntent
* Add config for single node runs * Fix notebook metadata * Run black formatter --------- Co-authored-by: Sharvin Jondhale <shjondhale@microsoft.com>
* [Fix]: Add accuracy also and reword informative assertions
* remove responsibleai-vision notebooks and related workflows
|
This pull request sets up GitHub code scanning for this repository. Once the scans have completed and the checks have passed, the analysis results for this pull request branch will appear on this overview. Once you merge this pull request, the 'Security' tab will show more code scanning analysis results (for example, for the default branch). Depending on your configuration and choice of analysis tool, future pull requests will be annotated with code scanning analysis results. For more information about GitHub code scanning, check out the documentation. |
| from azure.identity import InteractiveBrowserCredential, DefaultAzureCredential | ||
|
|
||
| try: | ||
| credential = DefaultAzureCredential() |
Check failure
Code scanning / CodeQL
Detect unsafe use of DefaultAzureCredential in python application Error
Sheri Gilley (sdgilley)
left a comment
There was a problem hiding this comment.
these changes will not break the docs builds.
* update rai-tabular environment version in rai tabular notebooks
* Add enable_rbac_authorization parameter to Key Vault * Update Docker image version for deployments
* Update Docker image version in entry.spec.yaml * Update Docker image version in entry.spec.yaml * Update Docker image to use CUDA 13.1 * Update Docker image to version 5.0 with CUDA 13.1 * Update conda.yaml * Update entry.spec.yaml * Update Docker image version in entry.spec.yaml * Update entry.spec.yaml * Update entry.spec.yaml * Update entry.spec.yaml * Update conda.yaml * Update Docker image to openmpi5.0-ubuntu24.04 * Update Docker image version in entry.spec.yaml * Change Python version from 3.10 to 3.8.12 * Update Docker image version in entry.spec.yaml * Update Docker image version in entry.spec.yaml * Update Docker image version in entry.spec.yaml * Update Docker image version in entry.spec.yaml
Co-authored-by: Arun <arunsu@microsoft.com>
…score_data failure (#3888) The sklearn-1.5/labels/latest curated environment updated to MLflow 2.19+ around Dec 16 2025. MLflow 2.19 uses the logged-models API which is not supported by AzureML tracking server, breaking mlflow.autolog() in the train step and mlflow.sklearn.load_model() in the predict/score_data step. Pin mlflow<2.19 at runtime in both train.yml and predict.yml commands.
… predict_step failure (#3889) The sklearn-1.5/labels/latest curated environment updated to MLflow 2.19+ around Dec 16 2025. MLflow 2.19 uses the logged-models API which is not supported by AzureML tracking server, breaking mlflow.autolog() in the train step and mlflow.sklearn.load_model() in the predict step. Pin mlflow<2.19 at runtime in both train.yml and predict.yml commands.
* Update command to install dependencies and run training * Fix command syntax for pip installation * update protobuf installation to avoid dependency conflicts * Update AzureML environment reference in notebook * Update AzureML environment reference in notebook
…nt dependencies (#3902) * Update e2e-ml-workflow.ipynb * Update AzureML environment reference in notebook * Update AzureML environment reference in notebook * Remove environment variables from online deployment Removed environment variables from deployment configuration. * Add environment variables to deployment configuration * Update azureml-mlflow version in notebook * Update environment version to 0.1.5 in ML workflow * Update Docker image version in ML workflow * Update environment version to 0.1.6 * update logging and saving of model artifact via MLflow * Fix deployment: custom score.py compatible with mlflow 3.x Three issues fixed: 1. Auto-generated MLflow scoring script imports azureml.ai.monitoring which is missing from auto-built env. Fixed with explicit conda env. 2. The mlflow.pyfunc.scoring_server.infer_and_parse_data and predictions_to_json functions were removed in mlflow 3.x. Rewrote score.py to parse input_data (split-oriented DataFrame) with pandas and serialize predictions with numpy/json directly. 3. Fixed model path: AZUREML_MODEL_DIR points to registration root but MLflow artifacts (MLmodel, model.pkl) are in a subdirectory. The score.py walks the directory to find the MLmodel file. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * Update e2e-ml-workflow.ipynb * Fix formatting in e2e-ml-workflow.ipynb * Format extra_pip_requirements for better readability * Update environment version to 0.1.7 * Update e2e-ml-workflow.ipynb --------- Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
#3928) * fix * Fix AzureML deployment env deps for inference server * Update pipeline.ipynb * Update conda.yaml * Update requirements.txt
…e change (#3948) * Initial plan * Fix storage account key extraction in managed identity notebooks * Handle list_keys API shape differences in managed identity notebooks * Fix managed identity notebook key extraction --------- Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com>
* Refactor code structure for improved readability and maintainability * Remove sample data download step and add input image for binary payloads deployment
* Fix uai notebook: handle azure-mgmt-msi Identity schema change * updated * Pin azure-mgmt-msi<8 and azure-mgmt-authorization<5 in UAI notebook * UAI notebook: extract principal/client id via as_dict() to support new azure-mgmt-msi schema
#3984) * version upgarde * json correct
#3980) * upgrade * image upgrade * upgrade * upgrade * try method added * fix * try method * fix
…te (#4035) The --set-default flag on batch-deployment create fails with a RequestInvalid error because the API can no longer find 'deployment_name' on BatchEndpointDefaults when sent through that code path. Replace --set-default with a two-step approach: create the deployment, then explicitly update the endpoint with --set defaults.deployment_name to set the default deployment. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* Fix amlsecscan root cron entrypoint ownership * Support cgroup v2 in amlsecscan wrapper * Harden amlsecscan install path * Format amlsecscan scanner changes * Move amlsecscan scheduled entrypoint under etc * Make amlsecscan reinstalls failure-safe Preserve trusted installations during package setup and stage root-owned scanner files with rollback before switching the cron schedule. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> --------- Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
|
You are seeing this message because GitHub Code Scanning has recently been set up for this repository, or this pull request contains the workflow file for the Code Scanning tool. What Enabling Code Scanning Means:
For more information about GitHub Code Scanning, check out the documentation. |
…4077) * Add SDK sample for batch deployment from registry pipeline component * Update batch endpoints readme with registry pipeline deployment sample * Format registry batch deployment sample notebook for black * Add CI workflow for from-registry batch pipeline notebook * Add discoverability README for registry batch deployment sample * Fix typo in autogenerated comment * Improve discoverability README for registry batch deployment sample * Add required discoverability keywords to registry sample README * Refine documentation and component loading logic Updated markdown and code comments for clarity. Added checks for component name and version. * Format registry batch deployment notebook * Normalize notebook JSON for black check * fix black format * Update Python version to use variable for flexibility * add owners for notebook * Remove azure-ai-ml installation command Removed the pip install command for azure-ai-ml from the notebook. * Fix JSON formatting in sdk-deploy-and-test.ipynb Fixed JSON formatting issue by removing an extraneous comma. * Remove azure-ai-ml==1.32.0 pin cell from from-registry batch pipeline notebook The notebook should validate against the CI-provided azure-ai-ml (dev-requirements, or the freshly-built release wheel via setup.sh), not downgrade to a pinned old SDK. Matches the sibling pipeline_with_components_from_yaml notebook, which has no install cell. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 48bc14df-6a6c-4fb3-bb66-4b187cc8163e * Fix from-registry batch sample: unique endpoint name, drop unsupported ComponentDeployment.Enabled property The endpoint used a fixed name and set properties={'ComponentDeployment.Enabled': True}. On re-run, begin_create_or_update tried to update the existing endpoint's immutable properties (also a True vs true case mismatch), failing with EndpointPropertiesUpdateNotSupported. Match sibling notebooks (hello-batch, training-with-components): append a random suffix so a fresh endpoint is created each run, and drop the unnecessary property (PipelineComponentBatchDeployment works without it). Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 48bc14df-6a6c-4fb3-bb66-4b187cc8163e --------- Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 48bc14df-6a6c-4fb3-bb66-4b187cc8163e
* Add batch endpoint job list workarounds to all batch endpoint notebooks Added alternative methods to list batch endpoint jobs when the standard approach encounters issues: Option 1: Using MLflow to search for runs by experiment name (same as the batch endpoint name) Option 2: Using ml_client.jobs.list() and filtering by experiment_name References: - ICM 569867940 - ICM 632272288 Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * Convert workaround code cells to markdown to prevent CI execution The workaround code cells were being executed by papermill during CI, causing ModuleNotFoundError for mlflow. Convert to a single markdown cell with fenced code blocks so the samples are visible to customers but not executed during CI runs. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * Clarify snippets assume default experiment name in batch endpoint notebooks Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: eb4ba443-58fa-4e89-89e0-d77baa4b8c18 --------- Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: eb4ba443-58fa-4e89-89e0-d77baa4b8c18
* feat: add AML command job migration sample Add an analyzer-first sample and CLI for migrating supported Azure Machine Learning command jobs to Microsoft Foundry Jobs, including copy and zero-copy data modes, RBAC preflight, documentation, and tests. Authored-by: GitHub Copilot for VS Code 0.60.2026073105 Model: GitHub Copilot (copilot) Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * Fix AML-to-Foundry migration safety checks * Fix migration sample formatting --------- Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
) * chore: upgrade workflow and dev Python dependencies * Update share-data-using-registry.ipynb * Update share-models-components-environments.ipynb * Update text-summarization-batch.ipynb * Update batch_driver.py * Update batch_driver.py * Update conda.yaml * Update imagenet-classifier-batch.ipynb * Update imagenet-classifier-batch.ipynb * Update dependencies and import statements in notebook * Change import from sys to os and update path * Update explore-data.ipynb * Update e2e-object-classification-distributed-pytorch.ipynb * Update explore-data.ipynb * style: format imagenet batch deployment files * fix * fix: stabilize notebook CI for spark/automl and update env pins * chore: default notebook workflow python fallback to 3.13
…_keras_minist_convnet-image_classification_keras_minist_convnet (#4108) * Update Docker image version in score.yaml * Update Docker image version for training component * Update Docker image version in prep_component.py --------- Co-authored-by: Alvin Ashcraft <73072+alvinashcraft@users.noreply.github.com>
…t storage account name (#4121) * Fix managed_vnet Spark example: use stable subscription-scoped default storage account name The managed_vnet Spark standalone example hardcoded a globally-unique default storage account name (sparkdefaultvnet) and only created it when 'az storage account check-name' reported the name as globally available. Storage account names are globally unique. Once 'sparkdefaultvnet' is held anywhere in Azure, check-name returns nameAvailable=false, so creation is skipped -- but the name is still substituted into the notebook, which then references a non-existent account in the ephemeral run resource group. The workspace deployment fails with: NotFound: Storage account .../storageAccounts/sparkdefaultvnet not found. Fix: - Derive the default storage account name from the subscription GUID (sparkdefvnet<first-8-hex-of-sub>). This is deterministic across runs (so AML-managed system private endpoints remain valid) and globally collision-proof (cannot be squatted by another tenant; always re-creatable in this subscription). - Create the account scoped to THIS resource group and gate only on whether it already exists here (idempotent), instead of gating on global name availability. Fail loudly on creation error rather than silently continuing into a confusing downstream workspace-create failure. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * Fix attached-Synapse Spark examples: provision Synapse in a SQL-capable region and bump Spark 3.3->3.5 The submit_spark_standalone_jobs and submit_spark_pipeline_jobs examples failed with 'Unknown compute target myattachedspark'. Root cause (from the setup step's bash -x log): 'az synapse workspace create' failed in eastus with: (SqlServerRegionDoesNotAllowProvisioning) Location 'eastus' is not accepting creation of new Windows Azure SQL Database servers for the subscription at this time. A Synapse workspace needs an underlying Azure SQL server; the eastus/subscription combo rejects new SQL servers, so the workspace was never created. That cascaded: spark pool create -> workspace ResourceNotFound; attach_managed_spark_pools.py -> 'Cannot find attached synapse workspace'; the AML compute 'myattachedspark' never existed; the notebook then died at its first job cell. The setup step's continue-on-error: true hid the real error. Changes (setup_spark.sh, else/attached-Synapse branch only): - Add SYNAPSE_LOCATION (default westus2) and provision the Synapse workspace + its ADLS Gen2 storage there, instead of the SQL-restricted eastus. - Give the Synapse Gen2 storage a region-distinct name (gen2ws2) so a fresh account is created in the new region. Storage accounts can't move regions, and the shared automation RG already holds an eastus-region '<rg>gen2' account; Synapse requires its primary storage in the workspace's region. - Bump the Synapse Spark pool from spark-version 3.3 (retired Mar 2025) to 3.5 (only currently supported Synapse runtime). - Quote the SQL admin password ('\'); it is generated from [:graph:] and was passed unquoted, so shell-special characters could break 'az synapse workspace create'. Note: the AML workspace is in eastus while the attached Synapse pool is now in westus2. If cross-region attach is not supported, the durable alternative is to drop the attached-Synapse cells and use serverless Spark only. Region may need CI iteration if westus2 is also SQL-restricted for the subscription. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * Try centralus for Synapse/SQL provisioning in Spark examples westus2 was also blocked by SqlServerRegionDoesNotAllowProvisioning for the CI subscription, matching the earlier eastus failure. Move the Synapse workspace and its co-located ADLS Gen2 storage to centralus (fresh gen2cus account, since storage cannot change region). Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * Make attached-Synapse setup idempotent and wait for provisioning centralus cleared the SqlServerRegionDoesNotAllowProvisioning block, so the Synapse workspace now provisions. The standalone and pipeline Spark gates run concurrently and derive the same Synapse workspace name, so a parallel run may already be provisioning it. Tolerate an existing workspace (WorkspaceNameUnavailable) and poll until the workspace reaches the Succeeded state before creating subresources, otherwise the Spark pool and firewall-rule calls fail with WorkspaceInInvalidStateForSubresourceCreateOrUpdate. Do the same for the Spark pool, and open the firewall before the data-plane role assignment so the runner IP is authorized (previously failed with ClientIpAddressNotAuthorized). Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> --------- Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Description
Checklist